
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Sowndappan S1 , Mythili S2 , Priya P3 , Dr. P. Sachidhanandam4 , Pavithra U5, Aarthi R S6
1Sowndappan S: Student, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India
2Mythili S: Student, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India
3Priya P: Assistant Professor, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India
4Dr. P. Sachidhanandam: Head of Department, Department of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India
5Pavithra U: Student, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India
6Aarthi R S: Student, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India
Abstract - Accurate object distance estimation using monocular cameras is challenging due to the absence of explicit depth information and the reliance on specialized hardwaresuchasstereovisionorLiDAR.Thispaperpresentsa lightweightcalibration-basedframeworkforestimatingobject distance using a single monocular camera. The proposed method uses camera height and tilt angle to construct a perspective grid that maps image coordinates to real-world grounddistanceintervals.Detectedobjectsarelocalizedusing a detection model, and the bottompixel ofeach bounding box is projected onto the calibrated grid to estimate distance. Unlike dense depth estimation methods, the proposed approach performs object-level distance estimation without requiring large training datasets or high computational resources. The method is object-agnostic and can be integrated with different detection models. Experimental resultsshowthatthesystemachieves100%intervalaccuracy and a mean absolute error of approximately 0.36 m over a distance range of 5 m to 15 m, excluding the near-field blind region. The system operates at approximately 5 frames per second on CPU-only hardware, demonstrating its suitability for real-time, low-cost monitoring applications.
Key Words: Monocular distance estimation, camera calibration, perspective grid mapping, object detection, geometric modeling, real-time systems, range-based estimation
Accuratedistanceestimationisafundamentalrequirement inmanycomputervisionapplications,includingsurveillance, autonomous navigation, robotics, and environmental monitoring.Whilehumansnaturallyperceivedepththrough binocular vision, enabling machines to estimate distances using visual input remains a challenging problem, particularly when relying on a single monocular camera. Unlike stereo vision systems, monocular setups do not provideexplicitdepthcues,makingdistanceestimationan inherentlyill-posedproblem.
Traditionalapproachesfordistanceestimationoftenrelyon specializedhardwaresuchasstereocameras,LiDARsensors, ordepthcameras.Althoughthesesystemscanachievehigh accuracy,theyintroducesignificantlimitationsintermsof cost, power consumption, and deployment complexity. In manyreal-worldscenarios suchaslarge-scalesurveillance systems or resource-constrained environments these requirementsmakesuchsolutionsimpractical.Asaresult, there is increasing interest in developing computationally efficient, monocular vision-based alternatives that can estimateobjectdistancewithoutadditionalhardware.
Recent advances in deep learning have led to significant progress in monocular depth estimation, where convolutionalneuralnetworksaretrainedtopredictdense depth maps from single images. Methods such as MonoDepth2 and DORN have demonstrated impressive performance on benchmark datasets. However, these approaches require large-scale annotated datasets, high computational resources, and often produce dense depth outputs that are unnecessary for applications focused on object-leveldistanceestimation.Moreover,theirdeployment on edgedevices orCPU-onlysystems remainschallenging duetotheircomputationalcomplexity.
To address these challenges, this paper proposes a calibration-based monocular object distance estimation frameworkthatcombinesgeometricmodelingwithobject detection.Theproposedmethodutilizescameraparameters suchasheightandtiltangletoconstructaperspectivegrid that maps image space to real-world ground distances Detectedobjectsarelocalizedusingadetectionmodel,and their positions are projected onto the calibrated grid to estimatetheirdistancefromthecamera.
Monoculardistanceestimationhasbeenextensivelystudied incomputervisionandcanbroadlybecategorizedintothree main approaches: deep learning-based depth estimation, geometry-based methods, and object-based distance estimationtechniques.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Recent advancements in deep learning have significantly improvedtheperformanceofmonoculardepthestimation. Thesemethodsaimtopredictdensedepthmapsfromsingle images using convolutional neural networks. Notable approaches such as MonoDepth2 utilize self-supervised learning techniques to estimate depth without requiring ground truth annotations, while DORN formulates depth estimation as an ordinal regression problem to improve accuracy.Althoughthesemodelsachievehighperformance onbenchmarkdatasetssuchasKITTIandNYUDepth,they present several limitations including large-scale dataset requirements, significant computational resources, and deploymentchallengesonresource-constraineddevices.
Geometry-basedapproachesrelyontheprinciplesofcamera projectionandcalibrationtoestimatedistances.Usingthe pinhole camera model, it is possible to relate image coordinates to real-world measurements when camera parameters such as focal length, height, and tilt angle are known.Thesemethodsarecomputationallyefficientanddo notrequiretrainingdata,makingthemsuitableforreal-time applications.However,manyexistingapproachesareeither limited to specific scenarios or lack robustness due to simplifiedassumptions.
Anothercategoryofmethodsfocusesonestimatingdistance usingobject-levelfeaturessuchasboundingboxsize,pixel location,orknownobjectdimensions.Whilethesemethods are simple and easy to implement, they often suffer from limited accuracy and poor generalization. Many rely on assumptionsaboutobjectsizeorrequirepriorknowledgeof objectdimensions,whichrestrictstheirapplicabilityacross differentobjectcategories.
From the above discussion, it is evident that existing approaches exhibit a trade-off between accuracy, computationalcomplexity,andpracticalapplicability.There is a clear need for a lightweight, calibration-based frameworkthatintegratesobjectdetectionwithgeometric modeling to provide accurate object-level distance estimation without requiring additional hardware or extensivetrainingdata.
Theproposedmethodaddressesthisgapbyintroducinga calibration-basedmonoculardistanceestimationframework that integrates object detection with perspective grid mapping. Unlike deep learning-based depth estimation
methods such as MonoDepth2 and DORN, the proposed approachdoesnotrequiredepthtrainingdataandoperates withsignificantlylowercomputationaloverhead.
Theproposedframeworkestimatesthedistanceofdetected objectsfromamonocularcamerausingacalibration-based geometricapproach.Thesystemintegratesobjectdetection withperspective-baseddistancemapping,enablingefficient object-leveldistanceestimationwithoutrequiringadditional sensorsordensedepthprediction.

Fig -1: Overall processing pipeline of the proposed system
Objects in the scene are localized using a bounding-boxbased detection model. For each detected object, the bounding box is defined as B = (xmin, xmax, ymin, ymax). The referencepointusedfordistanceestimationisthebottomcenter of the bounding box, defined as (xc, yb) = ((xmin + xmax)/2,ymax),wherexcrepresentsthehorizontalcenterand ybrepresentsthebottompixelcoordinate.Thebottompixel

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
is selected because it approximates the point of contact betweentheobjectandthegroundplane.
Theproposedmethodreliesonthefollowingknowncamera parameters:
• H:Cameraheightfromthegroundplane
• θ:Cameratiltanglerelativetothehorizontalaxis
• f:Effectivefocallengthofthecamera
• yc:Verticalcenteroftheimage
Themethodassumesaplanargroundsurface,fixedcamera positionandorientation,andthatobjectsareincontactwith thegroundplane.Theseassumptionssimplifythegeometric modeling and are valid for many surveillance and monitoringapplications.
Thecoreofthe proposedmethodisthetransformation of image coordinates into real-world ground distances using perspectivegeometry.Theverticalpositionofthedetected objectintheimageisconvertedintoanangulardeviation:α =arctan((yb -yc)/f).Usingcameraheightandtiltangle,the grounddistanceDiscomputedas:D=H/tan(θ+α).
To improve computational efficiency, the continuous distancespaceisdiscretizedintopredefinedintervalsusing a perspective grid. During runtime, the bottom pixel coordinateyb iscomparedagainstgridboundariesandthe correspondingdistanceintervalisassigned,enablingO(1) constant-timedistanceestimation.


The perspective grid boundaries are computed during an offlinecalibrationstage.CameraheightHandtiltangleθare measured,grounddistancesareselectedasreferencepoints, correspondingimagepositionsarecomputedviageometric projection,andboundaryvaluesarestoredinalookuptable. This design significantly reduces computational overhead andensuresconsistentperformanceacrossframes.
The proposed method achieves O(1) distance estimation complexityperobject.Itrequiresnodepthmodeltraining, introduces minimal runtime overhead beyond object detection, and enables real-time operation on CPU-only hardware
TheproposedsystemwasimplementedusingPythonwith OpenCVforvideoacquisitionandpreprocessing,Ultralytics YOLOv8 for object detection, and NumPy for numerical computations.Theimplementationwasdesignedtooperate efficientlyonCPU-onlyhardware.
Foreveryinputframe:(1)theimageispassedtotheobject detection module, (2) detected bounding boxes are extracted, (3) the bottom pixel of each bounding box is identified, (4) the pixel is mapped to the calibrated perspectivegrid,and(5)thecorrespondingdistanceinterval is assigned. This pipeline enables continuous monitoring withminimallatency.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
ObjectdetectionisperformedusingaYOLOv8-basedmodel trained on a custom dataset. The distance estimation componentremainsindependentofthedetectionmodeland can be integrated with any object detector. This modular design allows flexible deployment across different applicationdomains.
4.4
The complete system integrates object detection and geometric distance estimation into a unified real-time pipeline. The lightweight nature of the proposed method allows it to operate on CPU-only hardware without GPU acceleration.Thesystemincludesanalertmechanismthat notifies users when an object is detected along with its estimateddistance.
5.1
The proposed system was evaluated on a CPU-only setup consistingofanAMDRyzen75700Uprocessorwith8GB RAM.NodedicatedGPUwasusedduringtesting.Thesystem achieved an average inference latency of approximately 190msperframe(≈5fps).
A monocular camera with a resolution of 480×640 pixels was used. The camera was mounted at a height of 10 m abovethegroundandorientedatadownwardtiltangleof approximately45°.Underthisconfiguration,theobservable groundregionextendsfromapproximately5mto15m,with a near-field blind zone (0–5 m) due to field-of-view constraints.Thecameraconfigurationusedforcalibrationis illustratedinFig.4

Usingtheproposedcalibrationframework,theobservable region was discretized into the following predefined distanceintervals:
• 5–7m
• 7–9m
• 9–11m
• 11–13m
• 13–15m
Toevaluateperformance,50testsamplesweregenerated across the valid operating range (5–15 m). Evaluation metricsinclude:IntervalAccuracy(whetherthepredicted intervalcontainsthegroundtruth)andMeanAbsoluteError (MAE = (1/n)Σ| Dactual Dpredicted|), where Dpredicted is the midpointoftheassignedinterval.
Theobjectdetectionmoduleprovidesspatiallocalizationof objectswithinthescene.Detectedboundingboxesareused to extract the bottom pixel coordinate required for geometricmapping. Theproposedframework isdetectoragnostic and can be integrated with any object detection model.

-5. Sample system output showing detected object and estimated distance range overlaid on the camera frame.

International
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net
The proposed method was evaluated across distances ranging from 5 m to 15 m, excluding the near-field blind region.Atotalof50samplesweretestedacrossalldefined intervals. Table 1 presents a representative subset of the results.
Table -1. Sample Distance Estimation Results
Lightweight , calibrationbased
The system operates at approximately 5 fps (190 ms per frame) on CPU-only hardware. The distance estimation moduleintroducesnegligiblecomputationaloverheaddueto its constant-time lookup-based implementation. The majority of computational cost is attributed to the object detectionstage.
6.6 Limitations
• Near-field blind region (0–5 m) due to camera configuration
• Discretizationerrorcausedbyfixedintervalwidth
• Assumptionofplanargroundsurface
Theresultsshowthatall50sampleswereassignedtothe correctdistanceinterval.Futureworkwillextendevaluation tobenchmarkdatasetssuchasKITTIandNYUDepth.
6.3 Evaluation Metrics
Interval Accuracy: Accuracy=CorrectPredictions/Total Samples=50/50=100%
Mean Absolute Error (MAE): MAE ≈ 0.36 m (over 50 samples)
6.4 Comparison with Existing Methods
Table -2. Comparison with Existing Distance Estimation Methods
• Performancedependencyonobjectdetectionquality
7. DISCUSSION
The experimental results demonstrate that the proposed calibration-based monocular framework provides reliable range-basedobjectdistanceestimationusingonlyasingle camera and minimal geometric parameters. The system achieves 100% interval accuracy across 50 test samples, indicating that the perspective grid mapping consistently assignsobjectstothecorrectdistanceranges.
The observed MAE of approximately 0.36 m reflects the discretizationofthedistancespaceintofixedintervals.Since each interval spans 2 m, the midpoint approximation introduces inherent estimation error even when the predicted interval is correct. Reducing the interval width would proportionally decrease the MAE, at the cost of increasedcalibrationcomplexity.
Akeystrengthoftheproposedmethodisitscomputational efficiency,operatingviaconstant-timelookupwithnegligible overheadbeyondobjectdetection,runningat5fpsonCPUonlyhardware.Thedistanceestimationframeworkisalso model-independent, relying solely on geometric

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
relationships without dependency on object-specific features.
However, several limitations must be acknowledged: the assumptionofaplanargroundsurface,thenear-fieldblind region (0–5 m), fixed distance interval discretization, controlledevaluationconditions,andpartialdependencyon detection quality. Despite these limitations, the proposed methodoffersastrongbalancebetweenaccuracy,efficiency, anddeployability.
This paper presented a calibration-based monocular framework for object distance estimation using a single camera.Theproposedmethodleveragescameraheightand tilt angle to construct a perspective grid that maps image coordinates to real-world ground distance intervals. By associating the bottom pixel of detected object bounding boxeswiththiscalibratedgrid,thesystemenablesefficient object-leveldistanceestimationwithoutrequiringadditional sensorsordensedepthprediction.
ExperimentalresultsdemonstratedanMAEof0.36mover distances from 5 m to 15 m with 100% interval accuracy, operating at approximately 5 fps on CPU-only hardware. Unlikedeeplearning-basedmethodssuchasMonoDepth2, theproposedapproachdoesnotrequirelarge-scaledepth datasets or GPU acceleration, offering a computationally efficientandscalablealternative.
Future work will focus on extending the framework to handle non-planar environments, improving robustness under varying lighting and occlusion conditions, and developing adaptive calibration techniques. Additionally, integratingmoreadvanceddetectionmodelsandevaluating across diverse real-world scenarios will further enhance applicability
[1]D.ScharsteinandR.Szeliski,"Ataxonomyandevaluation ofdensetwo-framestereocorrespondencealgorithms,"Int. J.Comput.Vision,vol.47,no.1–3,pp.7–42,Apr.2002.
[2]R.HartleyandA.Zisserman,MultipleViewGeometryin ComputerVision,2nded.Cambridge,U.K.:CambridgeUniv. Press,2004.
[3] J. Levinson et al., "Towards fully autonomous driving: Systems and algorithms," in Proc. IEEE Intell. Veh. Symp., Baden-Baden,Germany,Jun.2011,pp.163–168.
[4]J.Redmon,S. Divvala,R.Girshick,andA.Farhadi,"You onlylookonce:Unified,real-timeobjectdetection,"inProc. IEEE Conf. Comput. Vision Pattern Recognit. (CVPR), Las Vegas,NV,USA,Jun.2016,pp.779–788.
[5] C. Godard, O. Mac Aodha, and G. J. Brostow, "Unsupervisedmonoculardepthestimationwithleft-right consistency," in Proc. IEEE Conf. Comput. Vision Pattern Recognit.(CVPR),Honolulu,HI,USA,Jul.2017,pp.270–279.
[6] T. Zhou, M. Brown, N. Snavely, and D. G. Lowe, "Unsupervised learning of depth and ego-motion from video,"inProc.IEEEConf.Comput.VisionPatternRecognit. (CVPR),Honolulu,HI,USA,Jul.2017,pp.1851–1858.
[7]A.Masoumian,H.A.Rashwan,J.Cristiano,M.S.M.Asif, and D. Puig, "Monocular depth estimation using deep learning: A review," Sensors, vol. 22, no. 14, p. 5353, Jul. 2022.
[8]Y.Kuznietsov,J.Stückler,andB.Leibe,"Semi-supervised deeplearningformonoculardepthmapprediction,"inProc. IEEE Conf. Comput. Vision Pattern Recognit. (CVPR), Honolulu,HI,USA,Jul.2017,pp.6647–6655.
[9] Z. Zhang, "A flexible new technique for camera calibration,"IEEETrans.PatternAnal.Mach.Intell.,vol.22, no.11,pp.1330–1334,Nov.2000.
[10]S.Foix,G.Alenya,andC.Torras,"Lock-intime-of-flight (ToF)cameras:Asurvey,"IEEESensorsJ.,vol.11,no.9,pp. 1917–1926,Sep.2011.
[11] D. Eigen, C. Puhrsch, and R. Fergus, "Depth map prediction from a single image using a multi-scale deep network,"inProc.Adv.NeuralInf.Process.Syst.(NeurIPS), vol.27,Montreal,QC,Canada,Dec.2014.
[12] R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, "Towards robust monocular depth estimation: Mixingdatasetsforzero-shotcross-datasettransfer,"IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 3, pp. 1623–1637,Mar.2022.
[13]H.Liang,X.Zhang,Y.Wang,andJ.Li,"Self-supervised object distance estimation using a monocular camera," Sensors,vol.22,no.8,p.2936,Apr.2022.
[14] S. Lee, J. Kim, H. Yoon, J. Shin, and H. Kim, "Vehicle distanceestimationfromamonocularcameraforadvanced driver assistance systems," Symmetry, vol. 14, no. 12, p. 2657,Dec.2022.
[15] M. Rezaei and R. Klette, "Computer vision for driver assistance: Simultaneous traffic and driver monitoring," SpringerTractsAdv.Robot.,vol.145,2017.
[16] S. Gasparini, P. Sturm, and J. P. Barreto, "Plane-based calibrationofcentralcatadioptriccameras,"inProc.IEEEInt. Conf.Comput.Vision(ICCV),RiodeJaneiro,Brazil,Oct.2007, pp.1–8.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
[17] Y. Yu and K. Hasegawa, "Relative distance estimation using monocular camera," J. Adv. Comput. Intell. Intell. Informat.,vol.14,no.6,pp.714–721,2010.
[18] M. A. Raza, A. Khattak, and H. Ali, "Framework for estimating distance and dimension of pedestrians in realtime using monovision," Neurocomputing, vol. 275, pp. 2572–2585,Jan.2018.
[19] M. Habibi, M. M. Nikkhah, and A. Aghdam, "Distance estimationbetweenmovingobjectsusingmonocularvision," inAIPConf.Proc.,vol.2591,no.1,p.080019,2023.
[20]B.U.Toreyin,Y.Dedeoglu,U.Gudukbay,andA.E.Cetin, "Computervisionbasedmethodforreal-timefireandflame detection,"PatternRecognit.Lett.,vol.27,no.1,pp.49–58, Jan.2006.
[21]K.Muhammad,J.Ahmad,I.Mehmood,S.Rho,andS.W. Baik,"Convolutionalneuralnetworksbasedfiredetectionin surveillancevideos,"IEEEAccess,vol.6,pp.18174–18183, 2018.
[22] G. Xu, Y. Zhang, Q. Zhang, G. Lin, and J. Wang, "Deep domain adaptation based video smoke detection using syntheticsmokeimages,"FireSaf.J.,vol.93,pp.53–59,Oct. 2017.
[23]G.Jocher,A.Chaurasia,andJ.Qiu,"YOLObyUltralytics," Jan.2023.[Online].Available:https://github.com/ultralytics/u ltralytics
[24] F. Amzajerdian et al., "Lidar systems for precision navigation,"NASATech.Rep.,2011.