Skip to main content

Case-Based Reasoning for Vision Friend

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Case-Based Reasoning for Vision Friend

1Student, Computer Science and Engineering, Dr. D. Y. Patil Polytechnic, Kolhapur, India

2Student, Computer Science and Engineering, Dr. D. Y. Patil Polytechnic, Kolhapur, India

3Student, Computer Science and Engineering, Dr. D. Y. Patil Polytechnic, Kolhapur, India

4Student, Computer Science and Engineering, Dr. D. Y. Patil Polytechnic, Kolhapur, India

Abstract: In the modern digital era, visually impaired users encounter multiple challenges in performing daily activities independently. Tasks such as navigation, object identification, reading printed text, and recognizing people often require external assistance. Traditional assistive tools like white canes, basic voice assistants, or manual help are limited in intelligence, adaptability, and real-time responsiveness. This project proposes Vision Friend, an intelligent AI-powered mobile assistance system that integrates computer vision, speech recognition, natural language processing (NLP), and text-to-speech technologies to provide real-time, context-aware support. The system processes live camera feeds and voice inputs to detect objects, recognize faces, extract text from images, and deliver immediate audio feedback. Additionally, Vision Friend incorporates adaptive learning to improve accuracy over time by learning from repeated user interactions. The proposed solution reduces dependency on human assistance, enhances user confidence and independence,andoffersascalable,efficient,andpractical accessibility solution.

Keywords:Object Detection, FaceRecognition,Speech Recognition,TextOCR, ComputerVision,MobileAssistance System,AI-PoweredSupport, Accessibility Tools.

1. INTRODUCTION

The rapid growth of smartphones and AI technologies hasopenednewpossibilitiesforassistivesystemsaimed at improving the quality of life for visually impaired individuals. Despite these advancements, many existing solutions fail to provide comprehensive real-time assistance in dynamic environments. Users often struggle with identifying nearby objects, reading signboardsordocuments,recognizingfamiliarfaces,and navigatingunfamiliarsurroundings.Traditionalassistive systemsaremostlyaudio-basedorrelyonphysicaltools that offer limited functionality and adaptability. Such systems lack intelligence to understand context, handle complex environments, or learn from user behavior. To overcome these limitations, intelligent assistive systems that combine multiple AI techniques are essential. Vision Friend is designed as an intelligent mobile-based solution that integrates object detection, face recognition, speech recognition, and OCR into a single application. By processing visual and auditory inputs

simultaneously, the system delivers fast, accurate, and personalized assistance. Unrecognized scenarios are stored and refined through adaptive learning, enabling continuous improvement. This approach significantly reduces response time, minimizes external dependency, and improves overall system efficiency and user satisfaction.

2. LITERATURE SURVEY

Several studies have contributed to the development of assistivetechnologiesusingAIandcomputervision:

Object detection models like YOLO enable real-time detection with high speed and accuracy, making them suitableformobileapplications.

FaceNet and CNN-based face recognition systems provide reliable identification of known individuals, enhancingsocialinteractionforvisuallyimpairedusers.

Deeplearning-basedspeechrecognitionsystemssuchas Deep Speech enable hands-free interaction and reduce usereffort.

OCR engines like Tesseract allowextraction of text from images,makingprintedcontentaccessiblethroughaudio output.

Research on adaptive learning and similarity measures highlights the importance of systems that improve over timebylearningfromuserinteractions.

Thesestudiescollectivelyindicatethatamulti-modalAIbased system can significantly enhance accessibility whenintegratedefficientlyintomobileplatforms..

3. PROPOSED SYSTEM

The proposed system, Vision Friend, is an intelligent mobile assistance application that provides real-time support tovisuallyimpairedusers. Thesystemoperates through voice commands or camera activation and processesinputsusingmultipleAImodules.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Core Functionalities:

ObjectDetection:Identifiesnearbyobjectsandobstacles toassistinnavigationandenvironmentawareness.

Face Recognition: Recognizes known individuals to supportsocialinteractions.

Text OCR: Reads printed or handwritten text from imagesandconvertsitintospeech.

Speech Recognition & NLP: Interprets user voice commandsandextractscontextualmeaning.

Audio Feedback: Delivers processed information throughspeechsynthesis.

Adaptive Learning: Stores new scenarios and refines modelsforfutureinteractions.

The system continuously learns from user interactions, ensuringimprovedaccuracy,fasterresponse, andbetter personalizationovertime.

Vision Friend Processing Cycle (Expanded)

1. Detect: The camera captures live video frames, which areanalyzedusingobjectandtextdetectionmodels.

2. Recognize: The system identifies detected objects, faces, or interprets speech inputs for context understanding.

3. Respond: Based on processed data, the system providesimmediateaudiofeedbackoralerts.

4. Learn: New or unclear scenarios are stored in the offlinedatabaseandusedtorefinemodelsforfutureuse.

System Architecture Diagram – Detailed

Explanation The Vision Friend system architecture consistsofthefollowingcomponents:

1.Input Sources :

Smartphone Camera (Live Feed): Captures real-time imagesand videoforobjectdetection,facerecognition, andOCR.

Microphone (Voice Commands): Accepts user voice inputsforcommandsandinteraction.

2. User Interface:

Accessible UI: Designed with large buttons, voice-based navigation, and minimal visual dependency to ensure easeofuseforvisuallyimpairedusers

3. Vision Friend Mobile Application (Core Layer)

This is the central control unit that coordinates all processing modules and manages data flow between inputsandoutputs.

4. AI Processing & Core Modules

ObjectDetectionModule(YOLO):Detectsobjectsinrealtimeandidentifiesobstaclesorimportantitems.

Face Recognition Module (CNN/FaceNet): Matches detected faces with stored profiles to identify known people.

Text-to-Speech&OCRModule(TesseractOCR):Extracts textfromimagesandconvertsitintoaudiblespeech.

Speech Recognition Module: Converts voice commands intotextandpassesthemtoNLPforinterpretation.

5. Offline Model Database

Stores learned scenarios, recognized patterns, and userspecificdata.

Enables functionality even with limited or no internet connectivity.

6. Output & Feedback Layer

Speakers/Headphones:Provideclearaudiofeedbackand instructions.

Vibration Motor (Haptic Feedback): Offers tactile alerts forcriticalnotificationssuchasobstacles.

Thismodulararchitectureensuresscalability,reliability, andefficientperformanceinreal-worldenvironments.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

4. COMPARATIVE ANALYSIS

Parameter ExistingSystem ProposedSystem

SystemType Traditional/Audioonlyaid IntelligentAI-based mobileassistance

QueryHandling Manualorbasic voice SpeechRecognition andNLP-based processing

ResponseTime Slowtomoderate Fastandreal-time

SimilarScenario Handling Limited Efficientdetection andrecognition

LearningCapability Noself-learning Self-learningusing adaptivemodels

EffortRequired High(externalhelp) Low

Scalability Limited High

5. ADVANTAGES

• Providesfasterresponsebydetectingandrecognizing elementsinreal-time.

• Reducesdependencyonhumanassistanceorphysical aids.

• Ensuresconsistentandreliablefeedbackforsimilar scenarios.

• Handleslargevarietiesofenvironmentalinputs efficiently.

• Improvesoveralluserindependenceandqualityoflife.

6. CONCLUSION

Vision Friend is an intelligent AI-powered mobile assistancesystem designed toaddress the limitations of traditionalassistivetoolsforvisuallyimpairedusers.By integrating object detection, face recognition, speech recognition, OCR, and adaptive learning, the system deliversaccurate,fast,andcontext-awareassistance.The detect–recognize–respond–learn cycle ensures continuous improvement and personalization. Overall, Vision Friend reduces user effort, enhances independence, and improves quality of life. Its scalable and self-learning architecture makes it suitable for modern mobile environments and provides a strong foundation for future advancements in accessibilityfocusedAIsolutions.

7. REFERENCES

[1] YOLO:Real-TimeObjectDetection JosephRedmon, 2016.

[2] FaceNet: A Unified Embedding for Face Recognition FlorianSchroffetal.,2015.

[3] Deep Speech: Scaling up End-to-End Speech Recognition AwniHannunetal.,2014.

[4] TesseractOCREngine RaySmith,2007.

[5] Coarse-to-FineVisionSystems JingLuetal.,2015.

[6] Computer Vision: Algorithms and Applications RichardSzeliski,2010.

[7] S.K.PalandS.C.K.Shiu,“SoftComputing,”2004.

[8] J.Luetal.,“Knowledge-BasedSystems,”2015.

[9] I.Watson,“ApplyingAITechniques,”1999.

Turn static files into dynamic content formats.

Create a flipbook
Case-Based Reasoning for Vision Friend by IRJET Journal - Issuu