
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Mrs. M. Usha1, Kota Chandini Grace2, Lakkoju Ganga Bhavani3, Majji Indumathi4 , Mandela Sharmila5
1 Faculty, Dept of Information Technology and Computer Applications, Andhra University College of Engineering for Women, Andhra Pradesh, India
2-5 B. Tech, Final year Student, Andhra University College of Engineering for Women, Andhra Pradesh, India
Abstract - Growing dependence on recorded lectures and online video content has made it considerably harder for students to extract and retain key academic information without investing substantial time in manual note-taking. This paper presents a lecture summarization framework that converts video lectures into structured study material by combining Faster-Whisper for speech-to-text transcription with DistilBART for abstractive summarization. The developed system accepts uploaded video files or YouTube links and produces concise summaries, key concepts, descriptive questions, and multiple-choice quizzes within a single unified pipeline. An academic refinement module transforms informal spoken language into coherent formal text, improving readability for direct revision use. Experimental evaluation using ROUGEmetricsdemonstrateshighersummarizationquality and content retention than baseline approaches across diverse lecture topics. These findings confirm that integrating transcription, summarization, and question generation into one cohesive tool effectively reduces cognitiveoverloadandsupportsself-directedlearning.
Key Words: Artificial Intelligence, Lecture Summarization, Natural Language Processing, DistilBART, Speech Recognition, Faster-Whisper, Automatic Question Generation, Educational Technology.
Thewidespreadadoptionofdigitallearningplatformshas fundamentally changed how educational content is delivered and consumed. Students increasingly rely on recorded lectures and online videos to understand complex academic subjects. However, the increasing length and volume of such content make it difficult to efficiently extract key information and revise important concepts. Studies in natural language processing indicate that unstructured learning content can lead to reduced comprehension and cognitive overload, especially when learners depend on passive video consumption [6], [8]. Traditional approaches such as manual note-taking are time-consuming and may result in incomplete or inconsistent understanding of lecture material.
Recent advancements in artificial intelligence and deep learning have enabled significant improvements in text understanding and generation. Transformer-based architectures such as those proposed in [1] have revolutionized natural language processing by effectively capturing contextual relationships within large datasets. Models such as BERT and BART further enhance text representation and summarization capabilities, producing coherent and context-aware outputs [4], [7]. Additionally, speech recognition systems such as Whisper enable accurateconversionoflectureaudiointotext,formingthe foundation for automated content processing [3]. These developments provide strong support for building intelligentlecturesummarizationsystems.
To address these challenges, this paper proposes the Smart Academic Lecture Summarizer, an artificial intelligence-based system designed to convert lecture videos into structured academic study material. The system processes input in the form of video links or uploaded files and performs speech-to-text transcription followed by transformer-based summarization. Unlike conventionaltoolsthatfocusonlyontranscriptionorbasic summarization,thistoolintegratesmultiplefunctionalities including key concept extraction, descriptive question generation, and multiple-choice question creation to supportactivelearningandself-assessment.
ThesystemisdevelopedusingPythonforcoreprocessing and integrates deep learning models such as Faster Whisper for speech recognition and DistilBART for summarization.Anacademicrefinementmoduleenhances readability by converting informal spoken language into structured academic text. The generated summaries are evaluated using standard metrics such as ROUGE, which measure the quality and coverage of the summarized content [6]. These evaluation techniques ensure that the system maintains both accuracy and coherence in its outputs.
Despite advancements in summarization and speech processing, most existing systems are limited to performing individual tasks such as transcription or summarization. Many approaches fail to provide structured academic outputs or lack integrated selfassessment features. Research in abstractive

International
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
summarization and sequence-to-sequence models highlights the importance of combining multiple techniques for effective content understanding [13], [14]. Furthermore, modern language models have demonstrated the potential for generating high-quality educational content, supporting automated learning systems[15].
This project aims to overcome these limitations by developing a unified system that integrates transcription, summarization, and content generation into a single pipeline. The developed framework enhances learning efficiency, reduces information overload, and provides a practical solution for transforming lecture videos into structured and interactive study material suitable for moderneducationalenvironments.
Transformer Architectures and Language Understanding: The survey explores how deep learning has transformed natural language processing tasks such as text summarization, content understanding, and information extraction. Transformer-based architectures have become thefoundationofmodernNLPsystemsduetotheirability to capture contextual relationships within large textual data, enabling efficient processing of long documents and lecture transcripts [1]. Pre-trained models such as BERT further enhance language understanding by generating contextual embeddings, while Sentence-BERT improves semantic similarity analysis, enabling accurate identification of key topics and important segments in textualdata[5],[7].
Summarization Models: Abstractive summarization techniquesusingtransformer-basedmodelssuchasBART and PEGASUS have significantly improved the generation of coherent and context-aware summaries by understanding the semantic meaning of input text rather than relying on simple extraction methods [4], [12]. Advancedapproachessuchaspointer-generatornetworks andreinforcementlearning-basedmodelsfurtherenhance summarizationqualityby balancingcontentcoverageand readability[13],[14].
Speech Recognition: Speech recognition technologies play a critical role in processing lecture-based content, where models such as Whisper enable accurate conversion of spoken language into text, even in diverse and noisy environments, forming the foundation for automated lecture analysis [3]. Modern language models have also demonstrated strong capabilities in generating educational content, supporting automated learning systems[15].
Evaluation Metrics and Research Gaps: Evaluation of summarization systems is commonly performed using ROUGE metrics, which measure the overlap between generated summariesand referencetext, ensuring quality
andrelevanceofoutputs[6].Despitetheseadvancements, most existing systems focus on individual tasks such as transcription or summarization and lack integration of multiple functionalities within a unified framework. Additionally, many approaches do not address the transformation of informal spoken language into structured academic text, which is essential for effective learning. This project addresses these gaps by integrating speech recognition, transformer-based summarization, and content generation techniques into a comprehensive system for lecture analysis and academic content generation.
TheSmartAcademicLectureSummarizerisdesignedasa structured processing pipeline that converts lecture videos into organized academic content using artificial intelligence techniques. The pipeline operates in five stages: (1) media acquisition, (2) speech-to-text transcription, (3) text preprocessing, (4) abstractive summarization, and (5) content generation, ensuring that raw lecture input is transformed into meaningful and structured learning material. The workflow begins with input acquisition, where the user provides either a YouTube link or uploads a video file. The system extracts the audio component using a media processing tool such as yt-dlp and prepares it for further analysis. The extracted audio is then processed using an optimized speechrecognitionmodel,Faster-Whisper,whichconverts spokenlanguageintotextefficiently.
Thetranscriptionprocessgeneratestimestampedoutputs, enablingsynchronizationbetweentheoriginallectureand thegeneratedcontent. The obtained transcript undergoes preprocessing, including noise removal, normalization, andsegmentation,to eliminateunnecessaryelementsand improve data quality. The segmented transcript is then processed using a transformer-based summarization model, DistilBART, which generates concise and contextaware summaries. The system adopts an abstractive summarizationapproach,allowingittorewritecontentin astructuredandcoherentmannerratherthanperforming simpleextraction.
To enhance the quality of the output, the system incorporates an academic refinement module that transforms informal spoken language into formal academic text by removing filler words and standardizing terminology. Additionally, a content generation module identifies key concepts and produces definitions, descriptive questions, and multiple-choice questions to supportactivelearningandself-assessment.
Finally, the system evaluates the generated summaries using performance metrics such as ROUGE scores and a customcoveragemeasuretoassesscontentretentionand

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
summary quality. The processed results are presented through an interactive user interface, providing a structured and efficient learning experience while maintainingcomputationalefficiencyandscalability.

The Smart Academic Lecture Summarization System is implemented as an intelligent framework that processes lecturevideosandconvertsthemintostructuredacademic contentusingartificialintelligencetechniques.Thesystem integrates speech recognition, natural language processing, and content generation modules to ensure efficient extraction andunderstanding of lecture material. Theimplementationofthedevelopedframeworkisshown inFig2

Thesystemprocessesinput intheformofa YouTubelink or uploaded video file. The input video is first converted into an audio stream using a media processing module, which is then passed to the speech recognition system. The generated transcript is further processed through
segmentation and preprocessing stages to remove noise and timestamps, ensuring structured and meaningful textual data. The processed data distribution and segmentationflowarerepresentedinFig.3.
The architecture of the developed system consists of two primary models: Faster-Whisper for speech recognition and DistilBART for text summarization. The system processes lecture data step by step, converting raw audio into structured academic content including summaries, keyconcepts,andassessmentmaterial.
The Faster-Whisper model is used for automatic speech recognition and transcription of lecture audio. It is designed using a transformer-based encoder-decoder architecturethat efficiently convertsspeech into text. The Faster-WhisperarchitectureisshowninFig.3

Input Feature Extraction Layer: Converts raw audio signals into log-Mel spectrograms, capturing timefrequency characteristics required for processing speech data.
Encoder Layer: Appliesmultipleself-attentionlayersto process audio features and capture temporal dependenciesandcontextualinformation.
Multi-Head Attention Layer: Enables the model to focus on different parts of the audio simultaneously, improvingspeechunderstandingandaccuracy.
Decoder Layer: Generates text tokens sequentially usingcross-attentionwithencoderoutputs.
Output Layer: Producesthefinaltranscriptbyselecting tokensbasedonprobabilitydistributions.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
The Faster-Whisper model enables efficient handling of long lecture recordings with reduced latency and improvedtranscriptionaccuracy.
The DistilBART model is used for abstractive text summarization. It is a lightweight transformer-based model that generates concise and meaningful summaries byunderstandingcontextualrelationshipswithinthetext. TheDistilBARTarchitectureisshowninFig.4.

Input Embedding Layer: Converts textual input into vector representations using token embeddings and positionalencoding.
Encoder Layer: Processes the input text using selfattentionmechanismstocapturecontextualrelationships.
Multi-Head Attention Layer: Allowsthemodeltofocus on multiple parts of the text simultaneously, improving semanticunderstanding.
Decoder Layer: Generates summary text sequentially usingcross-attentionwithencoderoutputs.
Feedforward Layer: Enhances feature extraction and improvesrepresentationlearning.
Output Layer: Producesthefinalsummarizedtextusing probability-basedtokenselection.
The DistilBART model ensures that long lecture transcripts are converted into concise and structured summarieswhilemaintainingcontextualcoherence.
The system further incorporates multiple processing modules to enhance the quality and usability of the
generated content. An Academic Refinement module is used to convert informal spoken language into structured academic text by removing filler words and standardizing terminology. This improves readability and ensures that theoutputresemblesformalstudymaterial.
A content generation module is implemented to extract key concepts and generate descriptive questions and multiple-choice questions. Important keywords are identifiedusingpattern-basedtechniques,andmeaningful questions are generated to support active learning and self-assessment.
Thesystemincludesanintuitivewebinterfacethatallows users to input lecture videos and view processed results. Thefrontendisdesignedusingwebtechnologiestoensure smooth interaction and real-time feedback. The system generatesmultipleoutputs,includingtranscript,summary, key concepts, descriptive questions, quiz and multiplechoicequestions,performancemetrics,andshare options. These outputs are presented in a structured format, enabling users to easily understand and utilize the generatedcontent.
To evaluate the effectiveness of the system, ROUGE-1 and ROUGE-L metrics are used to measure the similarity between the generated summary and the original transcript.Inaddition,acoveragemetricisusedtoassess how effectively key concepts are retained in the summarized output. These evaluation techniques ensure that the system produces accurate, relevant, and highqualityacademiccontent.
The overall system integrates all components into a unified pipeline, ensuring smooth data flow from input to output.Theimplementation isdesignedtoensuresmooth system performance and usability, making it suitable for classroomandself-studyenvironments.
The Smart Academic Lecture Summarizer was evaluated using multiple lecture videos of varying durations and subject domains to analyze its overall performance. The evaluation focused on the system’s ability to generate structured academic summaries, preserve key concepts, and provide additional learning components such as descriptive questions and multiple-choice questions. The systemwastestedonbothshortandlonglecturevideosto ensureconsistencyacrossdifferentlevelsofcomplexity. The complete summarization pipeline was implemented using a transformer-based approach, which includes multipleprocessingstagescontributingtothefinaloutput:

Research
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
• Speech-to-Text Module: Converts lecture audio into timestamped text using an optimized Faster-Whisper transcription model.
• Segmentation Strategy: Divides long transcripts into smaller segments to ensure balanced and efficient processing.
• Transformer-based Summarization: Generates concise academic summaries using the DistilBART model.
• Content Generation Module: Produces definitions, descriptive questions, and multiple-choice questions from summarizedcontent.





-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072


Fig.10: Display of automatically generated possible exam questions based on the analyzed lecture content.

Fig.11: MCQ quiz interface presenting objective questions for testing user understanding
From the observations, it is evident that the system effectively generates structured academic summaries that capture the essential concepts of the lecture. The segmentation strategy ensures uniform coverage of the entire lecture, preventing information loss. The summarization model maintains coherence and logical flow,makingtheoutputsuitablefordirectacademicuse.
Evaluation Metrics and Formulation
To measure the performance of the summarization system, standard evaluation metrics were used. These metrics compare the generated academic summary with the original transcript and help assess how much informationhasbeenpreserved.
ROUGE-1 measures the overlap of individual words (unigrams) between the reference transcript and the generated summary. It reflects how effectively important keywordsareretained.

ROUGE-L evaluates similarity based on the longest commonsubsequence(LCS).Itcapturesthestructuraland sequentialsimilaritybetweenthetranscriptandsummary.

Where T isthetranscriptand S isthegeneratedsummary.
Coverage Metric:
A custom coverage metric is used to measure the percentageofuniquetranscriptwordsretainedinthefinal summary.

Where:


=setofwordsinthetranscript
=setofwordsinthesummary
The summarization model maintains coherence and logical flow, making the output suitable for direct use as study material. The performance of the system was quantitatively evaluated using ROUGE metrics and a

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
customcoveragemetric.Theresultsforindividual lecture inputsarepresentedinTable1.
Table-1 ROUGE and Coverage Scores for Lecture Videos

Fig.12: ROUGE Score Comparison for Individual Lecture Inputs
Fig.12 shows the ROUGE score comparison for multiple lectureinputs,highlightingtheperformanceofthesystem incapturingkeyconceptsfromthetranscript.
The ROUGE scores indicate that the system achieves a strong overlap with the original transcript, particularly in terms of keyword extraction and sentence structure preservation. The ROUGE-1 score reflects effective retention of important terms, while ROUGE-L demonstrates that the sequence and contextual flow of informationaremaintainedinthegeneratedsummaries.
The system performance was further analyzed across multiple lecture samples, and the average results are presentedinTable2.
Table-2 Average ROUGE and Coverage Scores Across Lecture Samples

Fig.13: Average ROUGE and Coverage Scores Across Lecture Samples
Fig.13representstheaverageROUGEandcoveragescores across multiple lecture samples, demonstrating the consistencyandstabilityofthesystem.
For the overall evaluation, the ROUGE-1 score is approximately 0.50, while ROUGE-L is around 0.46, and the coverage metric reaches approximately 0.64. These valuesindicatethatthesystemmaintainsastrongbalance betweencontentretentionandsummaryconciseness.The results confirm that the proposed summarizer is effective ingeneratingstructuredacademicsummariessuitable for educationalapplications.
In addition to summarization, the system successfully generated supplementary learning materials such as key concepts, descriptive questions, and multiple-choice questions. These outputs were relevant to the lecture content and improved the overall usability of the system forrevisionandself-assessment.
From the results, it can be concluded that the developed tool effectively transforms lecture videos into structured academic content. The integration of transcription, summarization, and content generation into a single pipeline improves learning efficiency and reduces the effort required for manual note-taking. The system performs consistently across different lecture types and provides a practical solution for modern educational needs.
The system achieved a ROUGE-1 score of 0.50 and 64% content coverage across tested lectures, demonstrating thatintegratingspeechrecognitionandtransformer-based summarization into a single pipeline effectively reduces the effort of manual note-taking. This work unifies transcription via Faster-Whisper, abstractive summarization via DistilBART, and structured content generation into one platform, converting unstructured

International
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
lecture recordings into organised academic material. The developed framework thereby reduces manual effort, saves revision time, and improves learning efficiency for studentsacrossdiverseeducationalsettings.
Thesystemaddresseschallengesintraditionalnote-taking and content revision by utilizing optimized deep learning models and structured processing techniques. The proportional segmentation strategy ensures complete coverage of lecture content, while the academic refinement module enhances readability and maintains a formal structure. The integration of multiple functionalities, including summary generation, key concept extraction, question generation, and performance evaluation,makesthesystemhighlyeffectiveforacademic usage. The intuitive web interface further ensures accessibility for students with minimal technical knowledge, enabling broader adoption across different educationalenvironments.
The system holds strong potential for further enhancement through the development of a dedicated mobile application for Android and iOS platforms, allowinguserstoaccesssummarizedcontentandlearning materials anytime and anywhere. Future improvements may include multilingual support to expand accessibility, as well as the integration of more advanced generative models to improve the quality and diversity of generated questions.
As AI capabilities continue to expand alongside advances ineducationaltechnology,thissummarizercancontribute to improved learning outcomes by enabling efficient knowledgeextractionandpersonalizedstudyexperiences.
The future scope of this work includes extending the system to support adaptive learning mechanisms and integration with online learning platforms, providing a scalable and intelligent solution for modern digital education.
[1]A.Vaswanietal.,“AttentionIsAllYouNeed,”Advances in Neural Information Processing Systems (NeurIPS), 2017.
[2] T. Wolf et al., “Transformers: State-of-the-Art Natural Language Processing,” Proceedings of EMNLP: System Demonstrations,2020.
[3]A.Radfordetal.,“RobustSpeechRecognitionviaLargeScale Weak Supervision,” arXiv preprint arXiv:2212.04356,2022.
[4] M. Lewis et al., “BART: Denoising Sequence-toSequence Pre-training for Natural Language Generation, Translation, and Comprehension,” Proceedings of ACL, 2020.
[5]N.ReimersandI.Gurevych,“Sentence-BERT:Sentence Embeddings using Siamese BERT-Networks,” Proceedings ofEMNLP,2019.
[6] K. Papineni et al., “ROUGE: A Package for Automatic Evaluation of Summaries,” Proceedings of ACL Workshop, 2004.
[7] J. Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” ProceedingsofNAACL,2019.
[8] D. Jurafsky and J. H. Martin, Speech and Language Processing,3rded.,2023.
[9] S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780,1997.
[10] C. Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” JournalofMachineLearningResearch,2020.
[11] P. Liu and X. Chen, “Abstractive Text Summarization UsingDeepLearningModels,”IEEEAccess,vol.8,2020.
[12]J.Zhanget al.,“PEGASUS:Pre-trainingwithExtracted Gap-Sentences for Abstractive Summarization,” ProceedingsofICML,2020.
[13] S. See, P. J. Liu, and C. D. Manning, “Get to the Point: Summarization with Pointer-Generator Networks,” ProceedingsofACL,2017.
[14]R.Paulus,C.Xiong,andR.Socher,“ADeepReinforced Model for Abstractive Summarization,” Proceedings of ICLR,2018.
[15] T. Brown et al., “Language Models are Few-Shot Learners,” Advances in Neural Information Processing Systems(NeurIPS),2020.

KotaChandiniGrace,Student, Andhra University College of EngineeringforWomen

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072



LakkojuGangaBhavani,Student
Andhra University College of EngineeringforWomen
Majji Indumathi,Student,Andhra University College of EngineeringforWomen

MandelaSharmila,Student,Andhra University College of EngineeringforWomen
