District-level rice yield forecasting in uttarakhand, India: A comparative study of regularization, kernel-based, tree-based and neural network models
Abstract
Accurate pre-harvest forecasting of rice yields is critical for food security and climate adaptation in Uttarakhand, a topographically complex Himalayan state. This study presents a rigorous comparative assessment of five machine learning algorithms-Random Forest (RF), Elastic Net (ELNET), Support Vector Machine (SVM), Extreme Gradient Boosting (XGBoost), and Artificial Neural Network (ANN) for kharif rice yield prediction across 13 districts. Models were calibrated and validated using a 25-year dataset (1998-2023) of district-level yields and monthly agrometeorological variables from the NASA Power website employing a chronological split to simulate real-world forecasting conditions. Performance was evaluated using R², RMSE, nRMSE, and mean bias error (MBE). Random Forest emerged as the most robust and generalizable model, consistently delivering superior validation performance across districts (R² = 0.55-0.89) with low error and minimal bias. XGBoost exhibited near-perfect calibration but was prone to overfitting, particularly in data-sparse high-altitude regions. SVM demonstrated competitive, district-specific utility, while ELNET provided stable, interpretable predictions. ANN performance was inconsistent, limited by the modest training sample size. The findings establish that no single model is universally optimal; rather, a district-specific selection strategy or hybrid ensemble approach is recommended for maximizing predictive accuracy in heterogeneous agroclimatic zones.
How this paper connects to the literature. Drag to explore, click any node to open that paper.
