Dry bean seed classification using data preprocessing techniques and distance-based classifiers
Abstract
The development of automatic systems for agricultural applications, particularly for seed classification, is a task of paramount importance to the agro-industrial sector. Accurate recognition of seed varieties supports varietal purity and scalable processing. As one of the world’s most significant crops, dry beans underscore the essential need for precise seed classification to safeguard crop quality and enhance agricultural productivity. Traditional methods for dry bean seed classification involve manual and semiautomated approaches, that are inadequate for scale processing, suffering from slow throughput and reliance on human interpretation. In this work, we proposed a pipeline that is centered on Weighted K-Nearest Neighbors (WKNN) for classifying seven dry-bean varieties. Firstly, we remove outliers per class and features via the IQR rule, rebalance classes, and standardize features with Z-score normalization. The results of this methodology show an improvement in the accuracy, precision, recall, and F1-score metrics for the classification of 7 different dry bean seeds in comparison with other classical machine learning and state-of-the-art approaches.
Commun. Math. Biol. Neurosci.
ISSN 2052-2541
Editorial Office: [email protected]
Copyright ©2025 CMBN
Communications in Mathematical Biology and Neuroscience