Skip to main content

Auralis: An AI-Driven Digital Twin-Based Virtual Personal Assistant for Meeting Intelligence and Pro

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Auralis: An AI-Driven Digital Twin-Based Virtual Personal Assistant for Meeting Intelligence and Productivity Optimization

Arpit Yadav1 , Siddhu Kumar2 , Ambika Rajpurohit3 , Harsh Yadav4 , Prof. Ashwini More5

1,2,3,4 Diploma in Computer Engineering, Thakur Polytechnic, Mumbai, India

5Professor, Department of Computer Engineering, Thakur Polytechnic, Mumbai, India

Abstract - Modern professionals often have to manage several digital interactions at once, such as meetings, emails, and task coordination. However, human attention is limited, whichcanleadtomissedinformationandlowerproductivity.

This paper introduces Auralis, an AI-driven Digital Twin-based Virtual Personal Assistant designed to help users by automating meeting insights and routine communication tasks.

Unlike traditional meeting tools that mostly act as passive recorders, Auralis actively processes real-time conversation data and generates meaningful responses based on context. The system uses large language models built on Transformer architectures, combined with retrieval-augmented generation techniques that help maintain semantic memory. Real-time communication is supported through modern web-based protocols,allowingthesystemtocaptureandanalyzemeeting data effectively.

Beyond participating in meetings, Auralis also summarizes discussions, extracts actionable items, and helps with email management.ByusingaDigitalTwinmodel,thesystemlearns from user behavior and offers context-aware assistance over time.

Auralis aims not to replace human involvement but to lessen the burden of repetitive tasks and enable more efficient participation across various workflows.

1. INTRODUCTION

Remote collaboration is now a key part of modern work processes, especially with the rise of digital communication platforms. Individuals often find themselves attending numerous meetings, managing emails, and coordinating tasks in different settings. This growingdemandcanoverwhelmourcapacity,resultingin lower productivity and mental fatigue. While tools like video conferencing platforms improve communication, theystillrequireconstanthumanpresence and manual effort. The idea of a Digital Twin was first introduced to create virtual representations of physical systems for monitoring and simulation [3]. Recently, this concept has expanded to model human behavior and interactionpatterns.

In this context, a Digital Twin can learn from user activities,preferences,andcommunicationstyles,allowing forpersonalizedandadaptivesupport.Auralisbuildsonthis ideabyservingasanAI-drivenDigitalTwin-basedVirtual Personal Assistant. Rather than being a passive tool, the system actively processes real- time data from meetings anduserinteractions.

It understands conversational context, generates relevant responses,andaidsindecision-making.Thisispossibledue to the use of Transformer-based language models [1] and retrieval-augmented generation techniques [2], which together support real-time reasoning and long-term contextualmemory.

In addition to tracking meetings, Auralis also helps with taskmanagementandautomatingemails,functioningasa complete productivity assistant. By continuously learning from user behavior, the system aims to lessen repetitive tasks and provide support tailored to various activities. The goal is not to replace human involvement but to improve efficiency by allowing smarter and more flexible interactionwithdigitalsystems.

2. EASE OF USE

Easeofusereferstohoweasilyuserscanunderstandand operate a system with minimal effort and technical knowledge.

In intelligent systems, usability is crucial for real-world adoption.

Evenadvancedtechnologiesmayfailiftheyarehardto useorrequirealotoflearning.ForAuralis,easeofuseisa keydesignfocussincetheintendedusersarenotexpectedto haveexpertiseinartificialintelligenceorsystemsetup.The systemmanagescomplextaskslikelanguageunderstanding, contextualreasoning,andmemoryretrievalinternally.Ituses Transformer-based models [1] and retrieval-augmented generationtechniques[2].Theseprocessesremainhidden from the user, making interaction simple and intuitive. Auralisoperates withminimaluserinput.Afteraninitialsetupphase,where usersprovidebasicpreferencesandbehavioralinputs,the systembeginstoworkasaDigitalTwin.Itcanassistwith meetings,emails,andtaskmanagement.Insteadofrequiring

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

constantinput,itsupportssemi-autonomousoperation.This reducestheneedforcontinuoususerinvolvementwhilestill allowingcontrolwhennecessary.Theinteractionlayeruses standardwebcommunicationtechnologies[4],whichhelpit integratesmoothlywithexistingplatforms.Fromtheuser’ s pointofview,mostactionsinvolvestraightforwardtriggers likeenablingfeaturesorgrantingaccesspermissions.The system handles underlying processes, including real-time data managementandAI-baseddecision-making, without exposinganytechnicalcomplexity.Byprioritizingsimplicity ininteractionandminimizingmanualconfiguration,Auralis strikes a balance between advanced functionality and usability. This approach ensures that the system can be easily adopted and used as a practical digital assistant, rather than a complicated technical solution.

3. RELATED WORK

Recentadvancementsinnaturallanguageprocessinghave been driven by Transformer-based language models. These models have improved how systems understand andgeneratetextthatmakessenseincontext[1].Building upon this, retrieval-augmented generation (RAG) approaches combine language models with external knowledge sources. This allows systems to create more accurateandcontext-richresponses[2].Thesetechniques form the basis for modern intelligent assistants that need both reasoning and memory skills. The idea of a Digital Twin was first introduced to model and simulate physical systems in a virtual space [3]. Recently, this concept has beenexpandedtorepresentuserbehaviorandinteraction patterns.Thischangeallowsthedevelopment of systems that can adjust to individual preferences and workflows. It has opened new opportunities for creating personalized, context-aware digital assistants. Real-time communication tools like WebRTC have madeitpossible for web applications to exchange data smoothly and with minimal delay. This makes them ideal for systems needing live interaction, such as meeting support tools [4]. Meanwhile, research on human–AI interaction emphasizes the need to design systems that are predictable,clear,andeasytouse. This is crucial for building user trust and encouraging adoption [5]. From a system perspective, ideas from autonomous agent theory outline how intelligent systemscanrepresentuserswhilestayingfocusedontheir goals [6]. Additionally, memory-based structures, such as memory networks, highlight the need to keep and use long-term context in conversational systems [7]. Current meeting-relatedsolutionsmainlyfocusonrecording, transcribing,andsummarizingdiscussions[8].

While these tools are helpful for analyzing meetings afterward, they do not engage or assist during live interactions. Auralis enhances these methods by merging

real-time processing, semantic memory, and Digital Twin modeling. This allows for active participationandcontextaware help during meetings, as well as additional support for productivity tasks like email management and task tracking.RelatedWorkInfluencing Fig -1:RelatedWorkOverview

4. SYSTEM ARCHITECTURE

TheAuralisframeworkisbasedonamodularandrealtime AI system that facilitates intelligent assistance through a Digital Twin-based representation of the user. The overall system architecture is based on a pipeline-basedapproach,whereseveralcomponentsof the system collectively contribute to data input processing, context analysis, and finally, response generation.Theoverallsystemarchitectureisbasedon ensuringlowlatencyandaligningitselfwiththeuser’s behaviorandcommunicationstyle.

The overall system is divided into seven primary components, namely, (1) Audio Capture Module, (2) Speech-to-Text Engine, (3) Context Analyzer, (4) PersonaEngine,(5)ResponseGenerator,(6)Memory Store, and (7) Communication Interface. Fig. 3 representstheoverallworkflowofthesystem.Firstly, theoverallworkflowofthesystemisinitiatedby the Audio Capture Module, which collects real-time audio from meetings and/or user interaction. The collectedaudioisthensentforfurtherprocessingby theSpeech-to-TextEngine,whichgeneratestextfrom thecollectedaudio.Thegeneratedtextisthenanalyzed bytheContextAnalyzercomponentofthesystem.

Toensurethattheresponseisalignedwiththeuser’s styleofcommunication,thePersonaEngineincludes user-specificpreferencesandpatterns.Thisallowsthe system to provide responses that are similar to the user’sstyleofresponseinaparticularsituation.Then, theResponseGeneratorusesthecontextandpersona information,inadditiontolanguagemodels,tocreatea responsethatissignificantandunderstandable. AnothersignificantaspectoftheAuralisarchitectureis the Memory Store. This allows the system to retain contextoveraperiodoftime,enablingittorecallpast interactions. This is useful for retrieval-augmented generation, making the response provided by the Auralis system more accurate. Lastly, the

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

CommunicationInterfaceisresponsibleforfacilitating interactioninreal-timewithexternalapplications.This allows the Auralis system to integrate easily with a meeting system and other applications for effective communication. All these aspects of the Auralis architecture make it a powerful and intelligent assistantthatcanbeusefulinavarietyofapplications.

4.1 Audio Capture Module

TheAudioCaptureModuleisresponsibleforcontinuously receiving the incoming audio stream from the meeting environment.Itinterfacesdirectlywiththecommunication platform and performs basic preprocessing such as noise filteringandbufferingtoensurestableinputfordownstream components.Themoduleoperatesinreal-timetominimize delaysandmaintainsmoothinteraction.

4.2 Speech-to-Text Engine

Upon completion of basic audio preprocessing, the processed audio stream is passed seamlessly to the Speech-to-Text(STT)Engine,whichconvertsspoken input into textual form. Since all higher-level processing depends on text, the accuracy of this componentiscritical.Thegeneratedtranscriptsare timestampedandforwardedtotheContextAnalyzer forfurtherinterpretation.

4.3 Context Analyzer

The Context Analyzer interprets the conversational flowofthemeeting,identifyingintent,trackingongoing topics, and evaluating the relevance of different speakers.Usingtransformer-basedlanguagemodelsto capture contextual relationships within the dialogue [1],itdetermineswhetheraresponseisrequiredorif informationshouldbestoredforfuturereference.

4.4 Persona Engine

To ensure that system responses remain consistent withtheuser,AuralisincludesaPersonaEnginethat models communication style, preferences, and behavioralpatterns.Thepersonaisinitializedduring thesetupphaseandrefinedovertimeusinginteraction history.Thisenablesthesystemtogenerateresponses that reflect the user ’ s tone and decision making approach,whichiscrucialformaintainingtrustin human–AIinteractions[5].

4.5 Response Generator

The Response Generator produces replies by combining the current context with relevant past

information.Itusesaretrieval-augmentedgeneration (RAG)approach[2],whereimportantdataisretrieved frommemoryandincorporatedintotheresponse.The underlying model is based on the Transformer architecture[1],enablingcoherentandcontext-aware languagegeneration

4.6 Memory Store

The Memory Store maintains both short-term and long-terminformation,includingconversationhistory, key decisions, and user preferences. Inspired by memory network architectures [7], this component enablesthesystemtoretaincontextacrossinteractions andcontinuallyimproveitsperformanceovertime.It also supports post-meeting review by providing structuredaccesstopastdiscussions.

4.7 Communication Interface

The final output is delivered back to the meeting environment through a real-time communication interface based on WebRTC technologies [4]. This enablesseamlessintegrationwithexistingplatforms and ensures low-latency interaction. Within the meeting, the Digital Twin appears as an active participantrepresentingtheuser. Overall, this modular design separates perception, reasoning,andcommunicationprocesses,allowingthe systemtooperateefficientlywhileremainingscalable andextensibleforfutureimprovements.

Fig -2:Systemarchitecture

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

5. METHODOLOGY

ThedevelopmentofAuralisfollowsapipeline-based approachdesignedtosupportreal-timeinteractionand intelligent assistance through a Digital Twin model. Themethodologyintegratesmultiplestages,including speechprocessing,contextualunderstanding,persona alignment,responsegeneration,andtaskmanagement, intoaunifiedworkflow.Theoveralldesignfocuseson maintaininglowlatencywhileensuringthatthesystem remains consistent with user behavior. The process begins with capturing audio input from the meeting environment,whichisthenconvertedintotextusinga speech-to-textengine.Thistextualdataservesasthe primary input for further analysis. The system then performs contextual understanding by analyzing the conversationtoidentifyintent,keytopics,andrelevant information. Transformer-based models are used at thisstagetocapturerelationshipswithinthedialogue and generate meaningful representations of the conversation[1].

Once the context is established, the system incorporates user-specific information through a personamodelingmechanism. Thisstepensuresthatanygeneratedresponsereflects theuser’scommunicationstyleandpreferences.The personaiscontinuouslyrefinedusingpastinteractions, allowingthesystemtoadaptovertime.

To generate responses, Auralis uses a retrievalaugmented generation (RAG) approach [2], where relevantinformationisretrievedfromamemorystore andcombinedwiththecurrentcontext.Thisenables the system to produce responses that are both contextuallyappropriateandconsistentwithprevious interactions. The generated output is then delivered back to the meeting through a real-time communication interface [4].In addition to real-time response generation, Auralis includes a task management component that identifies and extracts actionableitemsfromconversationsandemails.Using natural language processing techniques, the system detects tasks, deadlines, and responsibilities, and organizesthemintoastructuredformat.Thesetasks arestoredandcanbereviewedorupdatedbytheuser, enabling better tracking of commitments discussed duringmeetings.Thesystemalsomaintainsamemory component that stores conversation history, key decisions,extractedtasks,anduserpreferences.This allows Auralis to improve its performance over repeatedinteractionsandprovidemoreaccurateand

context-aware assistance. Overall, the methodology ensuresabalancebetweenreal-timeresponsiveness, contextualaccuracy,andadaptivebehavior.

5.1 System Workflow

TheoperationalworkflowofAuralisbeginswhenthe Digital Twin connects to the user’s working environment, including online meetings and communicationplatforms.Formeetingscenarios,the system joins through a WebRTC-based interface [4], where incoming audio streams are continuously capturedandbufferedbytheAudioCaptureModule. ThebufferedaudioisthenforwardedtotheSpeech-toTextenginefortranscription.

Onceconvertedintotext,theconversationisprocessed bytheContextAnalyzer,whichdeterminessemantic intent,topicrelevance,andconversationalflowusing Transformer-basedrepresentations[1].Basedonthis analysis, the system decides whether to generate a response, store the information, extract tasks, or remainsilent.

If a response is required, the Persona Engine conditions the system using user-specific behavioral parameters.TheResponseGeneratorthenproducesa context-aware reply using a retrieval-augmented generationpipeline[2].Thegeneratedoutputisfinally transmitted back into the meeting through the communicationlayer,enablingreal-timeparticipation.

Inadditiontomeetinginteraction,Auralisextendsits workflow to email and task management. Incoming emailsareanalyzedusingnaturallanguageprocessing techniques to identify important information, categorize content, and detect actionable items. Similarly,duringmeetings,thesystemextractstasks, deadlines,andresponsibilitiesfromconversations. Thesetasksarestructuredandstoredforuserreview, allowing efficient tracking of commitments across differentworkflows.

Fig -3:Auralismethodology

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

5.2 Persona Modeling

Auralis employs a structured persona modeling approach to maintain alignment with the user’ s communicationstyle.Duringtheinitializationphase, usersprovidepreferencedatasuchastone(formalor informal), response length, domain expertise, and meetingpriorities.

Theseattributesformthebasepersonaprofile.

Thepersonaiscontinuouslyrefinedusinginteraction historyandfeedback.Thisadaptiveapproachensures that system responses remain consistent with user expectationsandimprovepersonalizationovertime, following established human–AI interaction principles [5].

5.3 Retrieval-Augmented Response Generation

Toimproveresponsequalityandcontextualgrounding, AuralisusesaRetrieval-AugmentedGeneration(RAG) mechanism[2].Theprocessinvolvesthreemainsteps: 1)Formulatingaquerybasedonthecurrentcontext 2)Retrievingrelevantinformationfrommemory 3)GeneratingaresponseusingaTransformer-based model

This approach helps reduce incorrect or irrelevant outputsandensuresthatresponsesarealignedwith bothcurrentandpastinteractions[1].

5.4 Memory Management Strategy

Auralis maintains both short-term and long-term memorylayers.Short-termmemorystorestheongoing meeting or interaction context, while long-term memory maintains historical data such as previous conversations, decisions, extracted tasks, and user preferences.

The memory system is inspired by neural memory network concepts [7], enabling efficient retrieval of relevantinformationwhenneeded.Aftereachsession, importantdiscussionsaresummarizedusingmeeting summarization techniques [8] and stored for future reference. This allows the system to maintain continuity across meetings and communication channels.

5.5 Task Extraction and Management

Auralisincludesataskmanagementmechanismthat automatically identifies actionable items from both meetings and emails. Using natural language processing, the system detects tasks, deadlines, and assigned responsibilities within conversations and writtencommunication.

Theseextractedtasksareorganizedintoastructured format and stored in the system, allowing users to review,update,orprioritizethem.Thisfeaturereduces

the need for manual note-taking and ensures that importantactionsarenotmissed.

5.6 Autonomous Decision Policy

Auralis does not respond to every input. Instead, it followsadecisionpolicytodeterminewhenaresponse isappropriate.

Thesystemevaluatesfactorssuchas:

• Relevanceofthespeaker

• Directnessofthequery

• Importanceofthetopic

• User-definedpriorities

This selective response strategy helps maintain naturalinteractionbehavior andavoids unnecessary interruptions. The decision-making process is influencedbyagent-basedsystemprinciples,wherethe DigitalTwinoperatesasanautonomousassistant[6].

5.7 Implementation Environment

TheprototypeimplementationofAuralisisdesigned fordeploymentinacloud-enabledenvironmentwith real-time processing capabilities. The system integrates:

• Transformer-basedlanguagemodelsforreasoning andgeneration[1]

• Retrieval-augmented pipelines for contextual grounding[2]

• WebRTC-based communication for real-time interaction[4]

• Persistent memory storage for long-term context retention

The modular design ensures scalability and allows futureextensions,suchasmultimodalinputsincluding videoandscreencontext.

Overall,theproposedmethodologyenablesAuralisto function as an intelligent Digital Twin capable of supporting meetings, managing communication, and organizing tasks, while maintaining user-specific behaviorandinteractionquality.

6. RESULTS AND DISCUSSION

This section evaluates the performance of Auralis across multiple functionalities, including meeting participation,taskextraction,andemailassistance.The objective is to assess response relevance, latency, persona alignment, and overall effectiveness of the DigitalTwininreducinguserworkload.

6.1 Experimental Setup

Auraliswastestedincontrolledvirtualenvironments simulating real-world usage. These included online meetings conducted using WebRTC-based

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

communication[4],alongwithsampleemaildatasets andtask-orienteddiscussions.Meetingsessionsranged from 10 to 45 minutes and included common professional interactions such as status updates, question-and-answer sessions, and decision-making discussions.

The system utilized Transformer-based language models for reasoning [1] and a retrieval-augmented generation (RAG) pipeline [2] for context-aware responses. All interactions were logged, including responsetime,extractedtasks,andgeneratedoutputs forevaluation.

6.2

Evaluation Metrics

The following metrics were used to evaluate system performance:

• Response Relevance: accuracy and contextual correctnessofgeneratedresponses

• Latency: time taken to generate a response after detectingaquery

• Persona Consistency:alignmentwithuser-defined communicationstyle

• Participation Accuracy: correctness in deciding whentorespond

• Task Extraction Accuracy:correctnessofidentified tasksanddeadlines

• User Effort Reduction:reductioninmanualeffort for meetings and communication These metrics evaluate both technical performance and usability aspectsofthesystem[5].

6.3 Quantitative Results

Thesystemdemonstratedstableperformanceacross different scenarios. The average response latency ranged between 25 and 80 seconds, depending on modelloadandcontextcomplexity,whichisacceptable forsemi-autonomousmeetingassistance.

Response relevance remained high in structured discussions, especially when supported by memory retrieval. The RAG-based pipeline improved contextual grounding and reduced irrelevant responsescomparedtostandalonegeneration.

Task extraction achieved consistent results in both meeting transcripts and email inputs, successfully identifyingkeyactions,deadlines,andresponsibilities. Thepersonamodellingmechanismmaintainedastable communicationstyleacrosssessions.

The decision policy effectively reduced unnecessary responses, allowing the system to participate only when required, which improved overall interaction quality

6.4 Qualitative Analysis

User-levelobservationsindicatethatAuralisbehaves morelikeanintelligentassistantratherthanapassive tool. Keystrengthsobservedinclude:

• Context-awareparticipationduringmeetings

• Consistentandpersonalizedcommunicationstyle

• Automaticextractionoftasksfromconversationsand emails

• Ability to recall previous discussions through memory

Theintegrationofmeetingintelligencewithemailand task management significantly reduced manual workload,especiallyinscenariosinvolvingfollow-ups andactiontracking.

However, certain limitations were observed. In meetingswithheavycross-talkorpooraudioquality, transcription accuracy decreased, affecting downstreamprocessing.Additionally,highlytechnical or domain-specific queries sometimes require more specializedknowledgethanisavailableinthecurrent system.

6.5 Comparative Discussion

Comparedtotraditionalmeetingtoolsandassistants, Auralis provides a more comprehensive solution. Existing systems primarily focus on transcription or summarization [8], whereas Auralis actively participates in conversations, manages tasks, and assistswithcommunication.

The combination of Digital Twin modeling [3], Transformer-based reasoning [1], and RAG-based memory[2]enablesmoreadaptiveandpersonalized behavior. Unlike basic assistants, the system also incorporatesanautonomousdecisionpolicy[6],which preventsexcessiveorirrelevantresponses.

6.6 Limitations and Future Improvements

Despite promising results, several areas can be improved:

• Reducingresponselatencyforreal-timeinteraction

• Improving performance in noisy or overlapping speechconditions

• Enhancingdomain-specificknowledgethroughbetter dataintegration

• Expanding email automation with advanced classificationandresponsegeneration

• Strengthening long-term learning for improved personalizationFutureworkwillexploremultimodal inputs (audio, video, and screen context) and more efficient model deployment techniques to improve real-timeperformance.

Overall,theresultsindicatethatAuralisisapractical

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

andscalablesystemforintelligentassistance,capable of supporting meetings, managing communication, andorganizingtaskswhilereducingusereffort.

7. CONCLUSIONS

This paper presented Auralis, an AI-driven Digital Twinbased Virtual Personal Assistant designed to support modernprofessionalworkflows.Thesystemaddressesthe growingchallengeofmanagingmultipledigitalinteractions, including meetings, emails, and task coordination, by reducing the need for continuous human involvement. Unlike traditional tools that focus only on recording or summarization,Auralisprovidesactiveassistancethrough real-time participation, contextual understanding, and intelligent task management. The proposed system integrates Transformer-based language models [1], retrieval-augmented generation mechanisms [2], Digital Twin principles [3], and real-time communication technologies [4] to enable context-aware reasoning and adaptivebehavior.Inadditiontoparticipatinginmeetings, Auralis extends its functionality to extracting actionable tasks,managingcommunication,andmaintaininglong-term contextualmemory.Thiscombinationallowsthesystemto functionasa comprehensiveproductivityassistant rather thanasinglepurposetool.

ExperimentalevaluationdemonstratesthatAuraliscan generate contextually relevant responses, maintain consistentpersonaalignment,andeffectivelyextracttasks from both meetings and emails. Although the response latency varies depending on system load, the overall performance is suitable for semi-autonomous assistance. The integration of memory and decision-making mechanisms further improves interaction quality and reduces unnecessary system responses. Despite these promisingresults,certainlimitationsremain.

The system performance can be affected in scenarios involving overlapping speech or poor audio quality. Additionally, domain-specific queries may require more specializedknowledgesourcestoimproveaccuracy.Future work will focus on reducing response latency, improving robustness in complex meeting environments, enhancing emailautomationcapabilities,andincorporatingmultimodal inputssuchasvideoandscreencontext.

Overall,Auralisdemonstratesthefeasibilityofcombining DigitalTwinmodelingwithmodernAItechniquestocreate an intelligent assistant capable of supporting meetings, communication, and task management. The proposed approach highlights the potential of human–AI collaboration systems to improve productivity while reducingcognitiveloadinreal-worldenvironments.

ACKNOWLEDGEMENT

Theauthorswouldliketoexpresstheirsinceregratitude toProf.AshwiniMoreforhervaluableguidance,continuous support,andencouragementthroughoutthedevelopmentof thisproject.Herinsightsandsuggestionsplayedacrucial roleinshapingthedesignandimplementationofAuralis.

REFERENCES

[1] Vaswani et al., ”Attention Is All You Need,” in Advances in Neural Information Processing Systems(NeurIPS),2017.

[2] P. Lewis et al., ”Retrieval-Augmented Generation forKnowledge-IntensiveNLPTasks,”inAdvances in Neural Information Processing Systems (NeurIPS),2020.

[3] M. Grieves, ”Digital Twin: Manufacturing Excellence through Virtual Factory Replication,” 2017.

[4] H. Alvestrand, ”WebRTC: Real-Time CommunicationinBrowsers,”RFC8825,2021.

[5] S. Amershi et al., ”Guidelines for Human-AI Interaction,” in Proceedings of the ACM Conference on Human Factors in Computing Systems(CHI),2019.

[6] M. Wooldridge, An Introduction to MultiAgent Systems,2nded.Wiley,2009.

[7] J. Weston et al., ”Memory Networks,” arXiv preprintarXiv:1410.3916,2014.

[8] G. Shang et al., ”Unsupervised Abstractive Meeting Summarization,” in Proceedings of the Annual Meeting of the Association for ComputationalLinguistics(ACL),2018.

[9] M. Hoy, ”Alexa, Siri, Cortana, and More: An Introduction to Voice Assistants,” Medical ReferenceServicesQuarterly,vol.37,no.1,pp.81–88,2018.

[10]S. Kiritchenko et al., ”Email Classification Using NaturalLanguageProcessing,”inProceedingsofthe Conference on Empirical Methods in Natural LanguageProcessing(EMNLP),2018.

Turn static files into dynamic content formats.

Create a flipbook