Skip to main content

A Lightweight Machine Learning Approach for Crop Yield Prediction in India Using Histogram Gradient)

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

A Lightweight Machine Learning Approach for Crop Yield Prediction in India Using Histogram

Gradient Boosting Regressor (HGBR)

1Student, Dept. of Computer Science Engineering, Sphoorthy Engineering College, Telangana, India

2Asst. Professor, Dept. of Computer Science Engineering, Sphoorthy Engineering College, Telangana, India

Abstract - Agriculture in India isn’t a recent development; it boasts an history of over 4 millennia having sustained the food security and livelihoods of an entire civilization. Agriculture has undeniably shaped the Indian identity , forming a heritage that every Indian takes great pride in. With a noble occupation come noble problems; and agriculture, though a noble pursuit, have its own unique challenges.Inmodernera,those‘nobleproblems’manifest as a struggle between traditional wisdom and modern demand. For generations, our ancestors relied on timehonored methods studying astronomical alignments (Panchang) and local biodiversity to accurately predict crop yields and the arrival of the monsoon. Rapid climate change has silenced the ancestral cues once used to predict yields, leaving today’s farmers caught between unpredictable seasons and rising demands. We must now find modern solutions that protect both our food security and the ancient spirit of Indian agriculture. Among the many proposed studies available, we utilized HistGradient BoostingRegressor (HGBR) to predict crop yields. This approach is particularly effective as it provides high accuracy while requiring low computational resources, making it a real-world application. Ultimately we stand at a crossroads where the sanctity of soil meets the precision of the algorithm. By utilizing HGBR, we offer low-resource, high-impact solution to the modern volatility that threatens our heritage. We seek not to replace the wisdom of the past, but to arm it with the tools of the future ensuring that the hands that feeds the nation are guided by the same certainty they held for millennia. With a deep reverence for the Earth and the legacy it sustains, in the name of the Divine who possessed the authority of earth and heavens, let us explore the remainder of the study.

Key Words: Histogram-based Gradient Boosting, Ensemble learning, Feature Binning, Machine Learning, Decision Trees, Predictive Modeling.

1. INTRODUCTION

Agriculture remains the backbone of the Indian economy, it has played a remarkable role in employment and food production, yet traditional yield prediction methods are failing under the bane of climate volatility and other uncertainty2. To restore this predictability, machine

learning especially ensemble methods like Gradient Boosting hasemergedasapowerfultoolfordata-driven forecasting4 .

1.1 Background of Agriculture and

Prediction

Crop

Agriculture has been a core part of India’s economy for more than four millennia, contributing significantly to employmentandfoodproduction2.Cropyieldpredictionis important for planning agricultural activities, managing supply chains, and ensuring national food security. Traditionally, our farmers have relied on indigenous knowledge such as their personal experience, seasonal trends, and environmental factors such as rainfall and temperaturetoestimatecropyield5

However, climate change has disrupted these natural patterns, leading to increased uncertainty in agricultural outcomes1. Irregular rainfall, sudden temperature changes, and extreme weather events have reduced the reliance on these traditional methods, creating a need for more advanced and accurate methods of predicting crop yield3

1.2 Problem Statement

One of the major challenges in modern agriculture is the unpredictability of crop yield which can fluctuate due to climate conditions and environmental factors6 When predictions are inaccurate, farmers and policymakers tendstomakepoordecisionswhichcanleadstoeconomic lossesanddisruptioninfoodsupply 4

To address these issues, it’s crucial to develop a model that can analyze the complex agricultural data and produce reliable yield predictions3. Such a model should takelesscomputationalresourcessothatitcanbeapplied in real-world scenarios with limited computational capacity7

1.3 Objective of the Study

Themainobjectiveofthisstudyistodevelopacropyield prediction model using Histogram based Gradient Boosting Regression (HGBR). The study aims to leverage

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

the advantages of histogram-based learning to enhance computational efficiency while maintaining high predictionaccuracy7 .

Furthermore, the focus of the study is to evaluate the performance of the proposed model using standard evaluation metrics and comparing its effectiveness with traditionalmachinelearningapproaches6 .

1.4 Scope of the work

The primary focus of this study is to predict crop yield using a structured dataset consisting of compatible agricultural features such as environmental conditions, soilproperties,andclimatic factors5.Theproposed model is designed to handle tabular data efficiently and provide accuratepredictions.

This work is confined to the implementation and evaluation of the HGBR model for crop yield prediction and does not cover real-time deployment or integration with IoT-based agricultural systems, which can be consideredforfutureresearch4

2. Literature Review

Thisliteraturereviewexaminestheevolutionofcropyield prediction from traditional statistical methods to advanced machine learning approaches, with a specific focusonhistogram-basedgradientboostingtechniques.

2.1 Traditional Crop Yield Prediction Methods

Statistical Methods

Traditionalcropyieldforecastinghashistoricallyreliedon a range of methodologies, primarily encompassing process-based models, crop simulation models, and conventional statistical techniques. These methods often utilize historical data analysis where experts in agricultural economics and farm management have long dependedon historical yielddata and farming-associative economic factors to project future production2 In addition, regression-based modeling is commonly employed where statistical approaches typically assume specificfunctionalforms,probabilitydistributions,ordata smoothness to establish correlations between variables3 . Furthermore physiological simulation plays a crucial role, as process-based or semi-physical models simulate crop growth by examining physiological processes and their interaction with the environmental components such as theplant-soil-atmospheresystem6

Limitations

While foundational, these conventional approaches face significant challenges in the modern agricultural landscape. Traditional methods, such as reliance on historical averages, often fail to capture the dynamic behavior of environmental factors like soil content, humidity, and rainfall, leading to inaccurate predictions4 Moreover, these approaches frequently ignore critical considerationssuchassoilnutrientlevels,moisturelevels, andpreciseweatherpatterns,whichcanleadtoimproper crop selection and long-term soil degradation5 Additionally, conventional econometric and statistical models are often insufficient for precisely capturing the complex, nonlinear agricultural issues and spatial fieldlevel variability that modern machine learning can address3 .

2.2 Machine Learning Approaches

Random Forest

Random Forest (RF) has emerged as a highly effective supervisedlearningmodelforagriculturalapplications:

High Accuracy: In recent studies, Random Forest models have achieved remarkable predictive accuracy, with one curateddatasetreaching99.15%5

Robust Performance: It is frequently utilized for both cropyieldpredictionandcroprecommendationduetoits abilitytohandlemulti-dimensionaldata4

Linear Regression

Linear regression remains a fundamental tool in the machinelearningrepertoireforyieldforecasting:

Predictive Mechanism: It predicts a measurable response by assuming a linear relationship between variouspredictorsandtheresponsevariable3

Comparative Performance:Whilesimplerthanensemble methods, multiple linear regression has shown competitive results in specific case studies, maintaining low mean squared errors alongside more complex algorithms3

Gradient Boosting

Gradient boosting techniques, such as XGBoost, represent amoreadvancedtierofpredictivemodeling:

Error Minimization: These algorithms are designed to iteratively improve model performance, often yielding verylowmeansquarederrorsincropyieldtasks4

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Handling Complexity: They are particularly adept at managingtheintricate,nonlinearrelationshipsinherentin agricultural systems that traditional models struggle to process5

2.3 Histogram Based Gradient Boosting(HGBR)

LightGBM/HGBR

Concept

Histogram-based Gradient Boosting (including LightGBM and Histogram-based Gradient Boosting RegressionHGBR) represents an optimization of the standard gradient boosting process. Instead of finding the optimal split point by iterating through all possible values of a feature, these models bin continuous feature values into discrete intervals (histograms). This significantly reduces the number of split points the algorithm needs to evaluate7

The efficiency of histogram-based methods stems from severalkeyarchitecturaladvantages:

Reduced Computational Cost:Byusingdiscretebins,the complexity of finding the best split is reduced from O(data) to O(Bins), allowing for much faster training on largedatasets7 .

Memory Efficiency: Storing discrete bins requires significantly less memory than storing continuous floating-pointvaluesforeverydatapoint7

Improved Scalability: These optimizations make the models particularly suitable for the "big data" challenges now common in agriculture, where large volumes of remote sensing and meteorological data must be processed7 .

3. Methodology

This section outlines the systematic approach used to develop the crop yield prediction model, covering data acquisition, processing, and the implementation of the Histogram-basedGradientBoostingRegressor(HGBR).

3.1 Dataset Description

The success of machine learning in agriculture depends heavily on the quality of environmental and cultivation parameters4 .

Source: The dataset used in this study is sourced from Kagglecontaininghistoricalrecordsofcropperformance.

Features: The model utilizes several key agricultural predictors,including:

Meteorological Data: Annual rainfall, average temperature,andhumidity1 .

Soil Parameters: Soil type and nutrient content (Nitrogen,Phosphorous,Potassiumlevels)5

Crop Information: Crop type and the specific season of cultivation4

Target Variable: The yield is measured in tonnes per hectare(ton/ha)orsimilarunits6 .

3.2 Data Preprocessing

Raw agricultural data is often inconsistent and requires cleaning to ensure model reliability2. The following steps wereperformedusingPandasandNumpy:

Missing Values Handling: Any null entries in the dataset wereaddressedthroughmeanimputation to prevent bias intheboostingprocess5

Encoding Categorical Data: Categoricalvariablessuchas 'Crop Type' and 'State/Region' were transformed into numerical formats using techniques like One-Hot Encoding or Label Encoding to make them compatible withscikit-learnalgorithms4

Normalization: To ensure that features with large scales (likerainfall)donotovershadowsmallerfeatures(likesoil pH), data normalization was applied to bring all values intoastandardrange3

3.3 Proposed Model HGBR:

The core of this research is the Histogram-based Gradient Boosting Regressor (HGBR), a modern evolution of the standard Gradient Boosting Machine (GBM).

Gradient Boosting: This is an ensemble technique that builds models sequentially. Each new tree attempts to correctthe errors (residuals)made by the previous trees, eventually"boosting"theoverallaccuracy4 .

Histogram Binning: Unlike traditional GBMs that evaluateeverypossiblesplitpointforeveryfeature,HGBR groups continuous features into discrete integer-valued bins(histograms)7

Performs fast: By operating on these bins rather than individual data points, the algorithm drastically reduces the number of split points to consider. This reduces computational complexity from O(nsamples) to O(nbins), makingitsignificantlymoreefficientforlargeagricultural datasets7 .

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

3.4 Model Implementation

The model was developed in a Python environment, leveraginghigh-performancescientificlibraries.

Tools and Libraries:

Python: Theprimaryprogramminglanguageusedfordata manipulation.

Scikit-learn: Specifically the HistGradientBoostRegressor module, which is optimized for datasets with more than 10,000samples7

Matplotlib: Used for visualizing model performance and errordistribution.

Hyperparameters:

Toachievethebestfitandavoid overfitting,thefollowing hyperparametersweretuned:

Learning Rate: Controls the contribution of each tree to thefinalresult(e.g.,0.1).

Max Depth: Limits the number of nodes in each tree to preventthemodelfrombecomingoverlycomplex7

Max Iterations: The total number of boosting rounds (trees)tobeconstructed7 .

Evaluation Metrics:

The model's performance was validated using a train-test split (typically80/20)andevaluatedbasedon:

R² Score: To measure the variance explained by the model3

Mean Squared Error (MSE) & Mean Absolute Error (MAE): Toquantifytheaveragepredictionerror6

Mean Squared Log Error (MSLE): To penalize underpredictionsmoreheavily,whichiscriticalinfoodsecurity planning3 .

4. Results and Discussion

This section presents the empirical findings of the study, evaluating the efficiency of the Histogram-based Gradient Boosting Regressor (HGBR) in predicting crop yields acrosstheIndianlandscape.

4.1 Performance Metrics

The model was evaluated using standard regression metrics.Basedontheexperimentalruns,theHGBRmodel achieved a high degree of predictive accuracy. The performanceissummarizedbelow:

Mean Absolute Error (MAE): The model recorded a low MAE, indicating that on average, the predicted yield deviatesonlyslightlyfromtheactualrecordedvalues.

Root Mean Squared Error (RMSE): The RMSE was calculatedtoaccountforlargervariances.Sincethemodel utilizes a log-transformation (np.log1p) during training, it effectivelyhandlesoutliersinyielddata.

R² Score: The model achieved an R² score of approximately 0.71, suggesting that the model explains over71%ofthevarianceinthecropyielddataset.

4.2 Experimental Results

A comparative analysis was performed to benchmark the HGBR model against traditional algorithms. The results indicate that ensemble-based boosting significantly outperformssimplelinearmodels.

Table 1: Comparative Performance of Models

Based Gradient Boosting Regressor

4.3 Graphical Analysis

The model's performance was further validated through visualdiagnostics:

Predicted vs. Actual Yield: Thescatterplot(Fig-3)shows astronglinearalignmentalongthe45-degreeidentityline. This indicates that the HGBR model accurately captures thetrendforbothlow-yieldandhigh-yieldscenarios.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Fig - 1 ActualvsPredictedplot

Residual Analysis: The residual plot provides a diagnostic look at the error distribution of the HGBR model.Theplotdisplaysthedifferencebetweentheactual and predicted values (residuals) on the y-axis against the predictedyieldonthex-axis.

Fig-2 ResidualvsPredictedPlot

Geospatial Distribution: Using the folium library and geospatialdata,thestudymappedcropdistributionacross India. This visualization confirms that the model successfully integrates latitude, longitude, and regional datatoaccountforlocalizedyieldvariations.

Fig-3 CropYieldwithnecessaryGeospatialdata

4.4 Discussion

The experimental results confirm that the Histogrambased Gradient Boosting Regressor is highly effective forlarge-scaleagriculturaldatasetsinIndia.

1.Analysis of HGBR Model Efficacy: Unlike Random Forest, which builds deep trees that can sometimes overfit, HGBR uses shallow trees built sequentially. By binning the continuous environmental features (like rainfall and temperature) into discrete histograms, the model reduces noise and focuses on the most significant datapatterns.

2.Geospatial Sensitivity: The inclusion of "Geo features" (Latitude, Longitude, and Region) allowed the model to accountforIndia'sdiverseagro-climaticzones.Thefolium visualization highlights that yield patterns are not uniform;themodel successfullyadaptedto theseregional differences.

3.Efficiency: Implementation showed that HGBR significantly reduced training time compared to standard GradientBoosting,makingitaviablesolutionforreal-time agriculturaldecision-supportsystems.

5.

Advantages and limitations

The implementation of the Histogram-based Gradient Boosting Regressor (HGBR) for crop yield prediction provides several technical and operational insights. This section evaluates the practical advantages and inherent constraintsoftheproposedapproach.

5.1 Technical Advantages

The HGBR architecture offers significant benefits over traditional Gradient Boosting and Random Forest models, particularly in the context of large-scale agricultural data. One of the primary advantage of HGBR is its use of histogram-basedbinning.Bygroupingcontinuousfeatures into discrete bins, the complexity of finding optimal split points is reduced from O(nsamples) to O(nbins). This resulted in significantly lower training latency during experimentation compared to standard ensemble methods5 Asagriculturaldatasetsgrowwiththeinclusion of multi-year remote sensing and soil sensor data, scalabilitybecomescritical.TheHGBRmodelisspecifically designed to handle datasets exceeding 10,000 samples with ease, maintaining high performance without a linear increase in memory consumption2 . Furthermore, the model demonstrated robust generalization, achieving an R2 score of 0.71. By iteratively reducing the residuals throughsequentialboosting,themodelcapturedcomplex, non-linear relationships between climatic variables such

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

as rainfall, temperature and crop yield that simple linear modelsfailedtoidentify1 .

5.2 Limitations and Constraints

Despite its high performance, the HGBR model presents certain challenges that must be addressed during implementation. One major limitation is its sensitivity to hyperparameters. The accuracy of the HGBR model is highly dependent on its configuration. Small changes in the learning_rate or max_depth can lead to significantly differentoutcomes.Forinstance,alearningratethatistoo high may cause the model to converge prematurely on a suboptimal solution, whereas a rate that is too low requires an excessive number of iterations5 . Another important constraint is the requirement for rigorous tuning. Unlike Random Forest models, which is relatively robust "out of the box," HGBR necessitates extensive hyperparameter optimization. Achieving the results presented in this study required careful balancing of the number of iterations (max_iter) and the tree depth to prevent overfitting, particularly when dealing with noisy Geospatial data.Furthermore, the model exhibitsa strong data quality dependency While HGBR is efficient, it remainssensitivetothequalityoftheinputhistograms.If the data preprocessing stage such as handling missing values or log-transformation is not performed correctly, the binning process may lead to a loss of information, negativelyimpactingthefinalprediction.

6. Conclusion

This research developed a predictive framework for Indian agriculture by using the Histogram-based Gradient Boosting Regressor (HGBR). By incorporating environmental variables and Geospatial coordinates, the study captures the complex, non-linear dynamics of crop productionacrossvariousagro-climaticzones.

TheHGBRmodeliscomputationallyefficientandscalable, enabling rapid, data driven decision making. Unlike conventionalmodelsthatstrugglewithlarge-scaledataor regional variations, the HGBR approach promotes rapid, data-driven decision-making. This is important for optimizing resource allocation, improving food security planning, and mitigating the economic risks faced by farmers due to unpredictable climatic shifts. The final experimentalresultsdemonstratethemodel'srobustness, achieving a high R² score of 0.71 and a low Mean Absolute Error.Thestrongcorrelationbetweenpredicted and actual yields, supported by the stable error distribution in residual analysis, demonstrates that HGBR is a superior architectural choice for modern crop yield forecasting. This study provides a foundational blueprint for deploying real-time, localized agricultural support systemsacrosstheIndiansubcontinent.

7. Future Works

Building upon the efficiency of the HGBR model, future research will seek to deepen the digital vision of Indian agriculture by transitioning toward Deep Learning architectures, such as LSTMs, to better capture long-term climatic patterns. The goal is to evolve this study into a real-time prediction system, providing farmers with instantaneous, data-driven insights rather than static historical analysis. By integrating this intelligence with Internet of Things (IoT) sensors for smart farming, we aim to create a living bridge between the soil and the cloud.Thistechnicalevolutionwillensurethattheancient wisdom of the land is permanently fortified by a continuous stream of real-time data, allowing the noble occupationtothriveamidstachangingworld.

REFERENCES

[1] Potential Impacts of Future Climate Changes on Crop Productivity of Cereals and Legumes in Tamil Nadu, India:AMid-CenturyTimeSliceApproach

[2] TransitionofIndianAgriculturefromGloriousPastto ChallengingFuture:ASeriousConcern

[3] AdoptionofMachineLearningMethodsforCropYield Prediction-based Smart Agriculture and Sustainable Growth of Crop Yield Production – Case Study in Jordan

[4] Machine Learning Based Crop Prediction and RecommendationSystem

[5] CropYieldPredictionusingMachineLearning

[6] Evaluating crop yield prediction models in Illinois using aquacrop, semi-physical model and artificial neuralnetworks

[7] LightGBM: A Highly Efficient Gradient Boosting DecisionTree

Turn static files into dynamic content formats.

Create a flipbook
A Lightweight Machine Learning Approach for Crop Yield Prediction in India Using Histogram Gradient) by IRJET Journal - Issuu