Komparasi Model Machine Learning untuk Estimasi Yield Tebu Berbasis Indeks Vegetasi Sentinel-2 di Kabupaten Jember

Abstract

Accurate sugarcane yield estimation is essential for farmers, the sugar industry, and agricultural policymakers, yet conventional field survey methods remain limited in spatial coverage, timeliness, and operational cost. This study compare the performance of three machine learning algorithms, RF, SVM, and DT, in predicting sugarcane yield in Jember Regency using Sentinel-2 imagery, and identifies the most influential indices as predictors. Research data were collected from 38 sugarcane plots across four sub-districts over one growing season. Five vegetation indices (NDVI, GNDVI, NDII, NDRE, and SAVI) were extracted via GEE as maximum, median, and mean values, yielding 15 candidate predictors. Variable selection was conducted sequentially through Spearman Rank Correlation and three iterations of VIF testing, resulting in three final predictors: GNDVIMax, NDIIMax, and NDREMax. Model evaluation employed LOOCV with R2, RMSE, MAE, and MAPE. SVM with RBF kernel achieved the best performance (R2 = 0,2987; RMSE = 11,8591 ton/ha; MAE = 8,5377 ton/ha; MAPE = 6,85%), followed by RF (R2 = 0,1814), and DT (R2 = 0,0039). Feature importance analysis indetified NDREMax as the most dominant predictor (55,3%), followed by NDIIMax (31,6%) and GNDVIMax (13,,2%), consistent with Spearman korelation findings. The low R2 values across all models were attributed to limited sample size, single compostie values per growing season, including variety, and soil conditions, that also influence sugarcane productivity. This study provides the first scientific baseline for remote sensing and machine learning-based sugarcane yield estimation in Jember Regency.

Description

Finalisasi Rudy k

Citation

Endorsement

Review

Supplemented By

Referenced By