
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Sanika Pramod Shetti¹, Janhavi Ganesh Bhise², Anuja Dnyandev Suryawanshi³, Neha Yuvraj Powar4
1Student, Computer Science and Engineering, Dr. D. Y. Patil Polytechnic, Kolhapur, India
2Student, Computer Science and Engineering, Dr. D. Y. Patil Polytechnic, Kolhapur, India
3Student, Computer Science and Engineering, Dr. D. Y. Patil Polytechnic, Kolhapur, India
4Student, Computer Science and Engineering, Dr. D. Y. Patil Polytechnic, Kolhapur, India
Abstract: In the modern digital era, visually impaired users encounter multiple challenges in performing daily activities independently. Tasks such as navigation, object identification, reading printed text, and recognizing people often require external assistance. Traditional assistive tools like white canes, basic voice assistants, or manual help are limited in intelligence, adaptability, and real-time responsiveness. This project proposes Vision Friend, an intelligent AI-powered mobile assistance system that integrates computer vision, speech recognition, natural language processing (NLP), and text-to-speech technologies to provide real-time, context-aware support. The system processes live camera feeds and voice inputs to detect objects, recognize faces, extract text from images, and deliver immediate audio feedback. Additionally, Vision Friend incorporates adaptive learning to improve accuracy over time by learning from repeated user interactions. The proposed solution reduces dependency on human assistance, enhances user confidence and independence,andoffersascalable,efficient,andpractical accessibility solution.
Keywords:Object Detection, FaceRecognition,Speech Recognition,TextOCR, ComputerVision,MobileAssistance System,AI-PoweredSupport, Accessibility Tools.
The rapid growth of smartphones and AI technologies hasopenednewpossibilitiesforassistivesystemsaimed at improving the quality of life for visually impaired individuals. Despite these advancements, many existing solutions fail to provide comprehensive real-time assistance in dynamic environments. Users often struggle with identifying nearby objects, reading signboardsordocuments,recognizingfamiliarfaces,and navigatingunfamiliarsurroundings.Traditionalassistive systemsaremostlyaudio-basedorrelyonphysicaltools that offer limited functionality and adaptability. Such systems lack intelligence to understand context, handle complex environments, or learn from user behavior. To overcome these limitations, intelligent assistive systems that combine multiple AI techniques are essential. Vision Friend is designed as an intelligent mobile-based solution that integrates object detection, face recognition, speech recognition, and OCR into a single application. By processing visual and auditory inputs
simultaneously, the system delivers fast, accurate, and personalized assistance. Unrecognized scenarios are stored and refined through adaptive learning, enabling continuous improvement. This approach significantly reduces response time, minimizes external dependency, and improves overall system efficiency and user satisfaction.
Several studies have contributed to the development of assistivetechnologiesusingAIandcomputervision:
Object detection models like YOLO enable real-time detection with high speed and accuracy, making them suitableformobileapplications.
FaceNet and CNN-based face recognition systems provide reliable identification of known individuals, enhancingsocialinteractionforvisuallyimpairedusers.
Deeplearning-basedspeechrecognitionsystemssuchas Deep Speech enable hands-free interaction and reduce usereffort.
OCR engines like Tesseract allowextraction of text from images,makingprintedcontentaccessiblethroughaudio output.
Research on adaptive learning and similarity measures highlights the importance of systems that improve over timebylearningfromuserinteractions.
Thesestudiescollectivelyindicatethatamulti-modalAIbased system can significantly enhance accessibility whenintegratedefficientlyintomobileplatforms..
The proposed system, Vision Friend, is an intelligent mobile assistance application that provides real-time support tovisuallyimpairedusers. Thesystemoperates through voice commands or camera activation and processesinputsusingmultipleAImodules.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Core Functionalities:
ObjectDetection:Identifiesnearbyobjectsandobstacles toassistinnavigationandenvironmentawareness.
Face Recognition: Recognizes known individuals to supportsocialinteractions.
Text OCR: Reads printed or handwritten text from imagesandconvertsitintospeech.
Speech Recognition & NLP: Interprets user voice commandsandextractscontextualmeaning.
Audio Feedback: Delivers processed information throughspeechsynthesis.
Adaptive Learning: Stores new scenarios and refines modelsforfutureinteractions.
The system continuously learns from user interactions, ensuringimprovedaccuracy,fasterresponse, andbetter personalizationovertime.
1. Detect: The camera captures live video frames, which areanalyzedusingobjectandtextdetectionmodels.
2. Recognize: The system identifies detected objects, faces, or interprets speech inputs for context understanding.
3. Respond: Based on processed data, the system providesimmediateaudiofeedbackoralerts.
4. Learn: New or unclear scenarios are stored in the offlinedatabaseandusedtorefinemodelsforfutureuse.

System Architecture Diagram – Detailed
Explanation The Vision Friend system architecture consistsofthefollowingcomponents:
1.Input Sources :
Smartphone Camera (Live Feed): Captures real-time imagesand videoforobjectdetection,facerecognition, andOCR.
Microphone (Voice Commands): Accepts user voice inputsforcommandsandinteraction.
2. User Interface:
Accessible UI: Designed with large buttons, voice-based navigation, and minimal visual dependency to ensure easeofuseforvisuallyimpairedusers
3. Vision Friend Mobile Application (Core Layer)
This is the central control unit that coordinates all processing modules and manages data flow between inputsandoutputs.
4. AI Processing & Core Modules
ObjectDetectionModule(YOLO):Detectsobjectsinrealtimeandidentifiesobstaclesorimportantitems.
Face Recognition Module (CNN/FaceNet): Matches detected faces with stored profiles to identify known people.
Text-to-Speech&OCRModule(TesseractOCR):Extracts textfromimagesandconvertsitintoaudiblespeech.
Speech Recognition Module: Converts voice commands intotextandpassesthemtoNLPforinterpretation.
5. Offline Model Database
Stores learned scenarios, recognized patterns, and userspecificdata.
Enables functionality even with limited or no internet connectivity.
6. Output & Feedback Layer
Speakers/Headphones:Provideclearaudiofeedbackand instructions.
Vibration Motor (Haptic Feedback): Offers tactile alerts forcriticalnotificationssuchasobstacles.
Thismodulararchitectureensuresscalability,reliability, andefficientperformanceinreal-worldenvironments.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
4. COMPARATIVE ANALYSIS
Parameter ExistingSystem ProposedSystem
SystemType Traditional/Audioonlyaid IntelligentAI-based mobileassistance
QueryHandling Manualorbasic voice SpeechRecognition andNLP-based processing
ResponseTime Slowtomoderate Fastandreal-time
SimilarScenario Handling Limited Efficientdetection andrecognition
LearningCapability Noself-learning Self-learningusing adaptivemodels
EffortRequired High(externalhelp) Low
Scalability Limited High
5. ADVANTAGES
• Providesfasterresponsebydetectingandrecognizing elementsinreal-time.
• Reducesdependencyonhumanassistanceorphysical aids.
• Ensuresconsistentandreliablefeedbackforsimilar scenarios.
• Handleslargevarietiesofenvironmentalinputs efficiently.
• Improvesoveralluserindependenceandqualityoflife.
6. CONCLUSION
Vision Friend is an intelligent AI-powered mobile assistancesystem designed toaddress the limitations of traditionalassistivetoolsforvisuallyimpairedusers.By integrating object detection, face recognition, speech recognition, OCR, and adaptive learning, the system deliversaccurate,fast,andcontext-awareassistance.The detect–recognize–respond–learn cycle ensures continuous improvement and personalization. Overall, Vision Friend reduces user effort, enhances independence, and improves quality of life. Its scalable and self-learning architecture makes it suitable for modern mobile environments and provides a strong foundation for future advancements in accessibilityfocusedAIsolutions.
7. REFERENCES
[1] YOLO:Real-TimeObjectDetection JosephRedmon, 2016.
[2] FaceNet: A Unified Embedding for Face Recognition FlorianSchroffetal.,2015.
[3] Deep Speech: Scaling up End-to-End Speech Recognition AwniHannunetal.,2014.
[4] TesseractOCREngine RaySmith,2007.
[5] Coarse-to-FineVisionSystems JingLuetal.,2015.
[6] Computer Vision: Algorithms and Applications RichardSzeliski,2010.
[7] S.K.PalandS.C.K.Shiu,“SoftComputing,”2004.
[8] J.Luetal.,“Knowledge-BasedSystems,”2015.
[9] I.Watson,“ApplyingAITechniques,”1999.