Data-Driven Crop Yield Forecasting Using Regression, Classification Methods: A Systematic Approach
Main Article Content
Abstract
Accurate prediction of Crop Yield Forecasting (CYF) is essential for improving Agricultural productivity, optimizing resource utilization, and supporting data-driven decision-making in precision farming. Traditional yield estimation methods often fail to capture the complex nonlinear relationships among environmental, soil, and cultivation factors. This study proposes CYF-CSVR, CYF-WKNN a Machine Learning-based Paddy Yield Forecasting and Classification model to enhance prediction accuracy. The first analysis is on predicting the yield and the second analysis is to do classification on the attribute ‘Yield Status’ -a binomial attribute which is created at the time of implementation. The experimental dataset was obtained from the UCI Machine Learning Repository Paddy dataset, which contains 45 agricultural attributes related to soil properties, weather conditions, crop management, and growth parameters. The dataset was preprocessed through missing value handling, normalization, and feature selection to improve model efficiency. Performance evaluation was conducted using standard regression metrics such as Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Relative Error(RE), and Correlation (r). In addition for the method CYF-WKNN, the performance is assessed with the classification metrics, Accuracy and Kappa Statistics. The results from Rapidminer demonstrate that the proposed model CYF-CSVR and CYF-WKNN effectively gives strong generalization capability. The proposed approach contributes to precision agriculture by enabling early yield forecasting, better crop management strategies, and ultimately supporting sustainable paddy cultivation.