International Research Journal of Engineering and Technology (IRJET)
e-ISSN: 2395-0056
Volume: 12 Issue: 05 | May 2025
p-ISSN: 2395-0072
www.irjet.net
AI DESKTOP ASSISTANT Sanjivani B. Adsul1, Aditya Ghurye2, Mahesh Dakore3, Mrunali Dhoke4, Kunal Dagade5 , Garvit Khandelwal6 1Professor, Department Artificial Intelligence & Data Science, VIT Pune, Maharashtra, India
2,3,4,5,6Students, Department of Artificial Intelligence & Data Science, VIT Pune, Maharashtra, India
---------------------------------------------------------------------***---------------------------------------------------------------------
Abstract - This paper presents an advanced AI assistant
Early results from the development of this assistant show its ability to seamlessly combine voice input processing, system automation, and intelligent decision-making to enhance user experience. The project underscores the potential of integrating these technologies, providing a powerful tool for improving productivity, system management, and personal interactions across various domains, such as home automation and professional productivity tools.
system that integrates voice recognition, automation, and intelligent decision-making by utilizing a combination of Google Speech Recognition, pyautogui, and a large language model (LLM) via the g4f.py library. Traditional automation systems often lack intelligent interaction and context understanding, limiting their versatility. The proposed assistant overcomes these limitations by processing voice commands to perform tasks ranging from web automation (such as controlling YouTube and Google Chrome) to managing system functions (including battery monitoring and PC restart). By incorporating an LLM, the assistant is capable of engaging in dynamic, context-aware conversations, enabling it to execute more complex tasks like writing code and answering questions. This hybrid framework provides a powerful tool for users, offering a seamless, multi-functional assistant that enhances both productivity and user experience across various platforms.
In recent years, AI-driven assistants have evolved from simple task automation tools to intelligent systems capable of complex decision-making and adaptive learning. Unlike traditional automation solutions, which rely on predefined scripts, modern AI assistants leverage machine learning algorithms to enhance their capabilities over time. With advancements in speech recognition, natural language processing (NLP), and automation frameworks, these assistants can interact in a human-like manner while performing a diverse set of operations. The ability to understand user intent, context, and execute crossapplication workflows makes AI-driven assistants a crucial element in smart computing environments.
Key Words: AI, Automation, Assistant, LLM, Intelligent
1.INTRODUCTION
2. LITERATURE REVIEW
The growing complexity of tasks and interactions in the digital age requires advanced solutions for intelligent automation and voice-driven assistance. Traditional automation tools typically operate through simple rulebased systems, lacking dynamic, context-aware interactions that are essential for more versatile applications. This project introduces an advanced AI assistant system that integrates multiple cutting-edge technologies, including Google Speech Recognition, pyautogui for web and system automation, and a large language model (LLM) via the g4f.py library.
The development of AI-driven personal assistants has seen significant advancements, particularly in integrating speech recognition and automation technologies. Guan et al. [1] explored the integration of large language models (LLMs) for process automation in intelligent virtual assistants, showcasing their ability to handle complex tasks and improve user experiences. Sharma et al. [2] highlighted the implementation of a voice-activated assistant (VOICEWISE) using speech recognition and natural language processing, demonstrating its potential in executing various automation tasks effectively. Mekni [3] provided a comprehensive analysis of conversational agents and their role in virtual assistants, emphasizing their application in facilitating seamless human-computer interactions. Richards [4] detailed the practical implementation of Anthropic's Computer Use Feature, showcasing its capabilities in automating user interactions with graphical user interfaces. Zhang and Shilin [5] surveyed GUI agents powered by large language models, highlighting advancements in their ability to interact with software interfaces and execute tasks autonomously. Ma et al. [6] proposed CoCo-Agent, a cognitive MLLM agent designed for smartphone GUI
The primary objective of this project is to develop a multifunctional AI assistant capable of performing a wide array of tasks. These tasks range from controlling media and browser applications like YouTube and Google Chrome to managing system functions such as battery monitoring and restarting the PC. The architecture of the system involves key components: Google Speech Recognition to process voice commands, pyautogui to automate interactions with web applications, and LLM integration to enable dynamic, context-sensitive conversations and more complex operations such as code writing and answering queries
© 2025, IRJET
|
Impact Factor value: 8.315
|
ISO 9001:2008 Certified Journal
|
Page 451