Skip to main content

Monocular Object Distance Estimation Using Calibration-Based Perspective Mapping and Detection Integ

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

Monocular Object Distance Estimation Using Calibration-Based Perspective Mapping and Detection Integration

Sowndappan S1 , Mythili S2 , Priya P3 , Dr. P. Sachidhanandam4 , Pavithra U5, Aarthi R S6

1Sowndappan S: Student, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India

2Mythili S: Student, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India

3Priya P: Assistant Professor, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India

4Dr. P. Sachidhanandam: Head of Department, Department of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India

5Pavithra U: Student, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India

6Aarthi R S: Student, Dept. of Information Technology, Knowledge Institute of Technology, Tamil Nadu, India

Abstract - Accurate object distance estimation using monocular cameras is challenging due to the absence of explicit depth information and the reliance on specialized hardwaresuchasstereovisionorLiDAR.Thispaperpresentsa lightweightcalibration-basedframeworkforestimatingobject distance using a single monocular camera. The proposed method uses camera height and tilt angle to construct a perspective grid that maps image coordinates to real-world grounddistanceintervals.Detectedobjectsarelocalizedusing a detection model, and the bottompixel ofeach bounding box is projected onto the calibrated grid to estimate distance. Unlike dense depth estimation methods, the proposed approach performs object-level distance estimation without requiring large training datasets or high computational resources. The method is object-agnostic and can be integrated with different detection models. Experimental resultsshowthatthesystemachieves100%intervalaccuracy and a mean absolute error of approximately 0.36 m over a distance range of 5 m to 15 m, excluding the near-field blind region. The system operates at approximately 5 frames per second on CPU-only hardware, demonstrating its suitability for real-time, low-cost monitoring applications.

Key Words: Monocular distance estimation, camera calibration, perspective grid mapping, object detection, geometric modeling, real-time systems, range-based estimation

1. INTRODUCTION

Accuratedistanceestimationisafundamentalrequirement inmanycomputervisionapplications,includingsurveillance, autonomous navigation, robotics, and environmental monitoring.Whilehumansnaturallyperceivedepththrough binocular vision, enabling machines to estimate distances using visual input remains a challenging problem, particularly when relying on a single monocular camera. Unlike stereo vision systems, monocular setups do not provideexplicitdepthcues,makingdistanceestimationan inherentlyill-posedproblem.

Traditionalapproachesfordistanceestimationoftenrelyon specializedhardwaresuchasstereocameras,LiDARsensors, ordepthcameras.Althoughthesesystemscanachievehigh accuracy,theyintroducesignificantlimitationsintermsof cost, power consumption, and deployment complexity. In manyreal-worldscenarios suchaslarge-scalesurveillance systems or resource-constrained environments these requirementsmakesuchsolutionsimpractical.Asaresult, there is increasing interest in developing computationally efficient, monocular vision-based alternatives that can estimateobjectdistancewithoutadditionalhardware.

Recent advances in deep learning have led to significant progress in monocular depth estimation, where convolutionalneuralnetworksaretrainedtopredictdense depth maps from single images. Methods such as MonoDepth2 and DORN have demonstrated impressive performance on benchmark datasets. However, these approaches require large-scale annotated datasets, high computational resources, and often produce dense depth outputs that are unnecessary for applications focused on object-leveldistanceestimation.Moreover,theirdeployment on edgedevices orCPU-onlysystems remainschallenging duetotheircomputationalcomplexity.

To address these challenges, this paper proposes a calibration-based monocular object distance estimation frameworkthatcombinesgeometricmodelingwithobject detection.Theproposedmethodutilizescameraparameters suchasheightandtiltangletoconstructaperspectivegrid that maps image space to real-world ground distances Detectedobjectsarelocalizedusingadetectionmodel,and their positions are projected onto the calibrated grid to estimatetheirdistancefromthecamera.

2. RELATED WORK

Monoculardistanceestimationhasbeenextensivelystudied incomputervisionandcanbroadlybecategorizedintothree main approaches: deep learning-based depth estimation, geometry-based methods, and object-based distance estimationtechniques.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

2.1 Deep Learning-Based Monocular Depth Estimation

Recent advancements in deep learning have significantly improvedtheperformanceofmonoculardepthestimation. Thesemethodsaimtopredictdensedepthmapsfromsingle images using convolutional neural networks. Notable approaches such as MonoDepth2 utilize self-supervised learning techniques to estimate depth without requiring ground truth annotations, while DORN formulates depth estimation as an ordinal regression problem to improve accuracy.Althoughthesemodelsachievehighperformance onbenchmarkdatasetssuchasKITTIandNYUDepth,they present several limitations including large-scale dataset requirements, significant computational resources, and deploymentchallengesonresource-constraineddevices.

2.2 Geometry-Based Distance Estimation

Geometry-basedapproachesrelyontheprinciplesofcamera projectionandcalibrationtoestimatedistances.Usingthe pinhole camera model, it is possible to relate image coordinates to real-world measurements when camera parameters such as focal length, height, and tilt angle are known.Thesemethodsarecomputationallyefficientanddo notrequiretrainingdata,makingthemsuitableforreal-time applications.However,manyexistingapproachesareeither limited to specific scenarios or lack robustness due to simplifiedassumptions.

2.3 Object-Based Distance Estimation Approaches

Anothercategoryofmethodsfocusesonestimatingdistance usingobject-levelfeaturessuchasboundingboxsize,pixel location,orknownobjectdimensions.Whilethesemethods are simple and easy to implement, they often suffer from limited accuracy and poor generalization. Many rely on assumptionsaboutobjectsizeorrequirepriorknowledgeof objectdimensions,whichrestrictstheirapplicabilityacross differentobjectcategories.

2.4 Research Gap

From the above discussion, it is evident that existing approaches exhibit a trade-off between accuracy, computationalcomplexity,andpracticalapplicability.There is a clear need for a lightweight, calibration-based frameworkthatintegratesobjectdetectionwithgeometric modeling to provide accurate object-level distance estimation without requiring additional hardware or extensivetrainingdata.

2.5 Positioning of the Proposed Work

Theproposedmethodaddressesthisgapbyintroducinga calibration-basedmonoculardistanceestimationframework that integrates object detection with perspective grid mapping. Unlike deep learning-based depth estimation

methods such as MonoDepth2 and DORN, the proposed approachdoesnotrequiredepthtrainingdataandoperates withsignificantlylowercomputationaloverhead.

3. PROPOSED METHODOLOGY

3.1 System Overview

Theproposedframeworkestimatesthedistanceofdetected objectsfromamonocularcamerausingacalibration-based geometricapproach.Thesystemintegratesobjectdetection withperspective-baseddistancemapping,enablingefficient object-leveldistanceestimationwithoutrequiringadditional sensorsordensedepthprediction.

Fig -1: Overall processing pipeline of the proposed system

3.2 Object Detection and Localization

Objects in the scene are localized using a bounding-boxbased detection model. For each detected object, the bounding box is defined as B = (xmin, xmax, ymin, ymax). The referencepointusedfordistanceestimationisthebottomcenter of the bounding box, defined as (xc, yb) = ((xmin + xmax)/2,ymax),wherexcrepresentsthehorizontalcenterand ybrepresentsthebottompixelcoordinate.Thebottompixel

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

is selected because it approximates the point of contact betweentheobjectandthegroundplane.

3.3 Camera Calibration and Assumptions

Theproposedmethodreliesonthefollowingknowncamera parameters:

• H:Cameraheightfromthegroundplane

• θ:Cameratiltanglerelativetothehorizontalaxis

• f:Effectivefocallengthofthecamera

• yc:Verticalcenteroftheimage

Themethodassumesaplanargroundsurface,fixedcamera positionandorientation,andthatobjectsareincontactwith thegroundplane.Theseassumptionssimplifythegeometric modeling and are valid for many surveillance and monitoringapplications.

3.4 Geometric Distance Estimation

Thecoreofthe proposedmethodisthetransformation of image coordinates into real-world ground distances using perspectivegeometry.Theverticalpositionofthedetected objectintheimageisconvertedintoanangulardeviation:α =arctan((yb -yc)/f).Usingcameraheightandtiltangle,the grounddistanceDiscomputedas:D=H/tan(θ+α).

To improve computational efficiency, the continuous distancespaceisdiscretizedintopredefinedintervalsusing a perspective grid. During runtime, the bottom pixel coordinateyb iscomparedagainstgridboundariesandthe correspondingdistanceintervalisassigned,enablingO(1) constant-timedistanceestimation.

3.5 Calibration Procedure and Lookup Table

The perspective grid boundaries are computed during an offlinecalibrationstage.CameraheightHandtiltangleθare measured,grounddistancesareselectedasreferencepoints, correspondingimagepositionsarecomputedviageometric projection,andboundaryvaluesarestoredinalookuptable. This design significantly reduces computational overhead andensuresconsistentperformanceacrossframes.

3.6 Computational Efficiency

The proposed method achieves O(1) distance estimation complexityperobject.Itrequiresnodepthmodeltraining, introduces minimal runtime overhead beyond object detection, and enables real-time operation on CPU-only hardware

4. SYSTEM IMPLEMENTATION

4.1 Implementation Environment

TheproposedsystemwasimplementedusingPythonwith OpenCVforvideoacquisitionandpreprocessing,Ultralytics YOLOv8 for object detection, and NumPy for numerical computations.Theimplementationwasdesignedtooperate efficientlyonCPU-onlyhardware.

4.2 Processing Pipeline

Foreveryinputframe:(1)theimageispassedtotheobject detection module, (2) detected bounding boxes are extracted, (3) the bottom pixel of each bounding box is identified, (4) the pixel is mapped to the calibrated perspectivegrid,and(5)thecorrespondingdistanceinterval is assigned. This pipeline enables continuous monitoring withminimallatency.

Fig -2: Geometric construction of perspective grid boundaries
Fig -3: Mapping from uniform image-space divisions to projected grid boundaries

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

4.3 Object Detection Integration

ObjectdetectionisperformedusingaYOLOv8-basedmodel trained on a custom dataset. The distance estimation componentremainsindependentofthedetectionmodeland can be integrated with any object detector. This modular design allows flexible deployment across different applicationdomains.

4.4

Real-Time System Integration

The complete system integrates object detection and geometric distance estimation into a unified real-time pipeline. The lightweight nature of the proposed method allows it to operate on CPU-only hardware without GPU acceleration.Thesystemincludesanalertmechanismthat notifies users when an object is detected along with its estimateddistance.

5. EXPERIMENTAL SETUP

5.1

Hardware Configuration

The proposed system was evaluated on a CPU-only setup consistingofanAMDRyzen75700Uprocessorwith8GB RAM.NodedicatedGPUwasusedduringtesting.Thesystem achieved an average inference latency of approximately 190msperframe(≈5fps).

5.2 Camera Configuration

A monocular camera with a resolution of 480×640 pixels was used. The camera was mounted at a height of 10 m abovethegroundandorientedatadownwardtiltangleof approximately45°.Underthisconfiguration,theobservable groundregionextendsfromapproximately5mto15m,with a near-field blind zone (0–5 m) due to field-of-view constraints.Thecameraconfigurationusedforcalibrationis illustratedinFig.4

5.3 Distance Segmentation

Usingtheproposedcalibrationframework,theobservable region was discretized into the following predefined distanceintervals:

• 5–7m

• 7–9m

• 9–11m

• 11–13m

• 13–15m

5.4 Evaluation Protocol

Toevaluateperformance,50testsamplesweregenerated across the valid operating range (5–15 m). Evaluation metricsinclude:IntervalAccuracy(whetherthepredicted intervalcontainsthegroundtruth)andMeanAbsoluteError (MAE = (1/n)Σ| Dactual Dpredicted|), where Dpredicted is the midpointoftheassignedinterval.

6. RESULTS AND EVALUATION

6.1 Object Detection Performance

Theobjectdetectionmoduleprovidesspatiallocalizationof objectswithinthescene.Detectedboundingboxesareused to extract the bottom pixel coordinate required for geometricmapping. Theproposedframework isdetectoragnostic and can be integrated with any object detection model.

-5. Sample system output showing detected object and estimated distance range overlaid on the camera frame.

Fig -4. Camera setup showing height and tilt angle used for calibration and distance estimation.
Fig

International

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net

6.2 Distance Estimation Results

The proposed method was evaluated across distances ranging from 5 m to 15 m, excluding the near-field blind region.Atotalof50samplesweretestedacrossalldefined intervals. Table 1 presents a representative subset of the results.

Table -1. Sample Distance Estimation Results

Lightweight , calibrationbased

6.5 Real-Time Performance

The system operates at approximately 5 fps (190 ms per frame) on CPU-only hardware. The distance estimation moduleintroducesnegligiblecomputationaloverheaddueto its constant-time lookup-based implementation. The majority of computational cost is attributed to the object detectionstage.

6.6 Limitations

• Near-field blind region (0–5 m) due to camera configuration

• Discretizationerrorcausedbyfixedintervalwidth

• Assumptionofplanargroundsurface

Theresultsshowthatall50sampleswereassignedtothe correctdistanceinterval.Futureworkwillextendevaluation tobenchmarkdatasetssuchasKITTIandNYUDepth.

6.3 Evaluation Metrics

Interval Accuracy: Accuracy=CorrectPredictions/Total Samples=50/50=100%

Mean Absolute Error (MAE): MAE ≈ 0.36 m (over 50 samples)

6.4 Comparison with Existing Methods

Table -2. Comparison with Existing Distance Estimation Methods

• Performancedependencyonobjectdetectionquality

7. DISCUSSION

The experimental results demonstrate that the proposed calibration-based monocular framework provides reliable range-basedobjectdistanceestimationusingonlyasingle camera and minimal geometric parameters. The system achieves 100% interval accuracy across 50 test samples, indicating that the perspective grid mapping consistently assignsobjectstothecorrectdistanceranges.

The observed MAE of approximately 0.36 m reflects the discretizationofthedistancespaceintofixedintervals.Since each interval spans 2 m, the midpoint approximation introduces inherent estimation error even when the predicted interval is correct. Reducing the interval width would proportionally decrease the MAE, at the cost of increasedcalibrationcomplexity.

Akeystrengthoftheproposedmethodisitscomputational efficiency,operatingviaconstant-timelookupwithnegligible overheadbeyondobjectdetection,runningat5fpsonCPUonlyhardware.Thedistanceestimationframeworkisalso model-independent, relying solely on geometric

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

relationships without dependency on object-specific features.

However, several limitations must be acknowledged: the assumptionofaplanargroundsurface,thenear-fieldblind region (0–5 m), fixed distance interval discretization, controlledevaluationconditions,andpartialdependencyon detection quality. Despite these limitations, the proposed methodoffersastrongbalancebetweenaccuracy,efficiency, anddeployability.

8. CONCLUSION

This paper presented a calibration-based monocular framework for object distance estimation using a single camera.Theproposedmethodleveragescameraheightand tilt angle to construct a perspective grid that maps image coordinates to real-world ground distance intervals. By associating the bottom pixel of detected object bounding boxeswiththiscalibratedgrid,thesystemenablesefficient object-leveldistanceestimationwithoutrequiringadditional sensorsordensedepthprediction.

ExperimentalresultsdemonstratedanMAEof0.36mover distances from 5 m to 15 m with 100% interval accuracy, operating at approximately 5 fps on CPU-only hardware. Unlikedeeplearning-basedmethodssuchasMonoDepth2, theproposedapproachdoesnotrequirelarge-scaledepth datasets or GPU acceleration, offering a computationally efficientandscalablealternative.

Future work will focus on extending the framework to handle non-planar environments, improving robustness under varying lighting and occlusion conditions, and developing adaptive calibration techniques. Additionally, integratingmoreadvanceddetectionmodelsandevaluating across diverse real-world scenarios will further enhance applicability

REFERENCES

[1]D.ScharsteinandR.Szeliski,"Ataxonomyandevaluation ofdensetwo-framestereocorrespondencealgorithms,"Int. J.Comput.Vision,vol.47,no.1–3,pp.7–42,Apr.2002.

[2]R.HartleyandA.Zisserman,MultipleViewGeometryin ComputerVision,2nded.Cambridge,U.K.:CambridgeUniv. Press,2004.

[3] J. Levinson et al., "Towards fully autonomous driving: Systems and algorithms," in Proc. IEEE Intell. Veh. Symp., Baden-Baden,Germany,Jun.2011,pp.163–168.

[4]J.Redmon,S. Divvala,R.Girshick,andA.Farhadi,"You onlylookonce:Unified,real-timeobjectdetection,"inProc. IEEE Conf. Comput. Vision Pattern Recognit. (CVPR), Las Vegas,NV,USA,Jun.2016,pp.779–788.

[5] C. Godard, O. Mac Aodha, and G. J. Brostow, "Unsupervisedmonoculardepthestimationwithleft-right consistency," in Proc. IEEE Conf. Comput. Vision Pattern Recognit.(CVPR),Honolulu,HI,USA,Jul.2017,pp.270–279.

[6] T. Zhou, M. Brown, N. Snavely, and D. G. Lowe, "Unsupervised learning of depth and ego-motion from video,"inProc.IEEEConf.Comput.VisionPatternRecognit. (CVPR),Honolulu,HI,USA,Jul.2017,pp.1851–1858.

[7]A.Masoumian,H.A.Rashwan,J.Cristiano,M.S.M.Asif, and D. Puig, "Monocular depth estimation using deep learning: A review," Sensors, vol. 22, no. 14, p. 5353, Jul. 2022.

[8]Y.Kuznietsov,J.Stückler,andB.Leibe,"Semi-supervised deeplearningformonoculardepthmapprediction,"inProc. IEEE Conf. Comput. Vision Pattern Recognit. (CVPR), Honolulu,HI,USA,Jul.2017,pp.6647–6655.

[9] Z. Zhang, "A flexible new technique for camera calibration,"IEEETrans.PatternAnal.Mach.Intell.,vol.22, no.11,pp.1330–1334,Nov.2000.

[10]S.Foix,G.Alenya,andC.Torras,"Lock-intime-of-flight (ToF)cameras:Asurvey,"IEEESensorsJ.,vol.11,no.9,pp. 1917–1926,Sep.2011.

[11] D. Eigen, C. Puhrsch, and R. Fergus, "Depth map prediction from a single image using a multi-scale deep network,"inProc.Adv.NeuralInf.Process.Syst.(NeurIPS), vol.27,Montreal,QC,Canada,Dec.2014.

[12] R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, "Towards robust monocular depth estimation: Mixingdatasetsforzero-shotcross-datasettransfer,"IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 3, pp. 1623–1637,Mar.2022.

[13]H.Liang,X.Zhang,Y.Wang,andJ.Li,"Self-supervised object distance estimation using a monocular camera," Sensors,vol.22,no.8,p.2936,Apr.2022.

[14] S. Lee, J. Kim, H. Yoon, J. Shin, and H. Kim, "Vehicle distanceestimationfromamonocularcameraforadvanced driver assistance systems," Symmetry, vol. 14, no. 12, p. 2657,Dec.2022.

[15] M. Rezaei and R. Klette, "Computer vision for driver assistance: Simultaneous traffic and driver monitoring," SpringerTractsAdv.Robot.,vol.145,2017.

[16] S. Gasparini, P. Sturm, and J. P. Barreto, "Plane-based calibrationofcentralcatadioptriccameras,"inProc.IEEEInt. Conf.Comput.Vision(ICCV),RiodeJaneiro,Brazil,Oct.2007, pp.1–8.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

[17] Y. Yu and K. Hasegawa, "Relative distance estimation using monocular camera," J. Adv. Comput. Intell. Intell. Informat.,vol.14,no.6,pp.714–721,2010.

[18] M. A. Raza, A. Khattak, and H. Ali, "Framework for estimating distance and dimension of pedestrians in realtime using monovision," Neurocomputing, vol. 275, pp. 2572–2585,Jan.2018.

[19] M. Habibi, M. M. Nikkhah, and A. Aghdam, "Distance estimationbetweenmovingobjectsusingmonocularvision," inAIPConf.Proc.,vol.2591,no.1,p.080019,2023.

[20]B.U.Toreyin,Y.Dedeoglu,U.Gudukbay,andA.E.Cetin, "Computervisionbasedmethodforreal-timefireandflame detection,"PatternRecognit.Lett.,vol.27,no.1,pp.49–58, Jan.2006.

[21]K.Muhammad,J.Ahmad,I.Mehmood,S.Rho,andS.W. Baik,"Convolutionalneuralnetworksbasedfiredetectionin surveillancevideos,"IEEEAccess,vol.6,pp.18174–18183, 2018.

[22] G. Xu, Y. Zhang, Q. Zhang, G. Lin, and J. Wang, "Deep domain adaptation based video smoke detection using syntheticsmokeimages,"FireSaf.J.,vol.93,pp.53–59,Oct. 2017.

[23]G.Jocher,A.Chaurasia,andJ.Qiu,"YOLObyUltralytics," Jan.2023.[Online].Available:https://github.com/ultralytics/u ltralytics

[24] F. Amzajerdian et al., "Lidar systems for precision navigation,"NASATech.Rep.,2011.

Turn static files into dynamic content formats.

Create a flipbook
Monocular Object Distance Estimation Using Calibration-Based Perspective Mapping and Detection Integ by IRJET Journal - Issuu